Strategy·14 min read

Long-Horizon AI Sales Agents Just Shipped. The Reliability Math Says Start Them Warm

Updately Team·2026-09-14

The outbound agent just got a clock measured in weeks

On 11 September 2026, Salesforce expanded Agentforce with seven job-ready agents and, buried under the naming exercise, shipped something more consequential than any of them: a long-horizon runtime. Hunter, the outbound sales agent, is the first agent to run on it. Salesforce describes Hunter as working "a sales pipeline from research to outreach, collaborating with sellers over weeks and months." It is in pilot now, with general availability in November 2026.

That is a real category shift for long-horizon AI sales agents, and it deserves to be understood precisely rather than celebrated or dismissed. Until now, every AI tool in your outbound stack was effectively a function call. You asked it for research, a message, a list, a summary. It returned an artifact. The unit of work was a request. What Salesforce shipped is an agent whose unit of work is an objective that persists across sessions, days, and quarters — "rescue my at-risk deals before end of quarter" — with the agent building the plan, choosing the tasks, and deciding which steps need your approval.

Here is the honest read for anyone running a sales team, an SDR function, or a GTM agency. The capability is real and the direction is correct. But the reliability research on long-horizon agents is unambiguous about one thing: success rates do not decline gently as horizons extend. They hold, then collapse. Which means the variable that decides whether a multi-week outbound agent produces pipeline or produces a mess is not the model. It is the quality of the goal you hand it on day one, and the freshness of the signal underneath that goal.

This post covers what actually changed, what the agent-reliability literature says about where long runs break, and a concrete way to pilot long-horizon outbound in the next six weeks without torching your domain, your LinkedIn account, or your brand.

What actually shipped on September 11

Seven agents, one that matters for outbound

Salesforce introduced Casey (support), Paige (IT and HR), Carter (shopper), Marshall (supply chain), Piper (inbound pipe gen), Fin (customer experience), and Hunter (outbound sales). Six are generally available now. Hunter is the pilot, GA November 2026, and the only one running on the new runtime.

The customer numbers Salesforce published are worth reading carefully, because they are the first concrete agentic-GTM figures the company has put in public. Perk reports 60% of its sales pipeline built by its outbound agent. Asana reports 4x conversation volume from Piper, its website agent, with an average 45-day deployment. Salesforce says it has delivered 7 billion Agentic Work Units across Agentforce and Slack, 3.2 billion of those in Q2 alone.

Those are adoption and volume metrics. They are not efficiency metrics, and they are not quality metrics. "60% of pipeline built by the agent" tells you nothing about the win rate on that pipeline versus human-sourced pipeline, which is the number a CRO actually needs. SiliconANGLE's coverage notes the same gap. Ask for it in your evaluation.

The three primitives underneath

The runtime itself is the interesting part, and Salesforce is explicit about the three capabilities that make it work:

  • Memory — context and progress persist across sessions, so the work does not reset when the conversation ends.
  • Durable execution — plans keep running over time; the agent can resume or course-correct when circumstances change.
  • Dynamic steering — behaviour adapts to an individual seller's feedback and direction.

PPC Land's write-up frames it as goals over weeks rather than chats, which is the right framing. Salesforce also shipped Agent Script, an open-source language for combining probabilistic AI reasoning with deterministic rules — a tacit admission that you do not want a model improvising every decision across a six-week run.

Note what those three primitives are, though. Memory, durable execution, and steering are exactly the three things the research community has identified as the failure surfaces of long-horizon agents. Salesforce did not solve them. It productised the places where they break.

Why "weeks, not chats" is genuinely hard

The research says performance collapses, it does not fade

The most useful recent work here is HORIZON, a cross-domain diagnostic benchmark from researchers at Wisconsin–Madison, Berkeley, and Georgia Tech, published in April 2026. They ran 3,100-plus agent trajectories across four domains (web, OS, embodied, database) using frontier GPT-5 and Claude models, systematically extending how many sequential subtasks each task required.

Three findings matter for anyone about to buy a multi-week outbound agent:

  1. Degradation is non-linear. Success rates stay relatively stable as you add subtasks, then drop sharply. There is no gentle slope warning you that you are approaching the limit.
  2. The failure composition changes, not just the failure rate. As horizons extend, planning failures (bad subplanning) and memory failures (catastrophic forgetting) become the dominant failure modes. Early-horizon failures are simple execution errors; late-horizon failures are the agent quietly working toward the wrong thing.
  3. Model choice stops mattering past the breaking point. In the web, OS, and database domains, performance gaps between models narrowed sharply once tasks entered the breaking region — success rates converged toward low values regardless of which frontier model was driving.

That third point is the one to sit with. The vendor conversation you are about to have in November will be about model quality. The research says that past a certain horizon length, model quality is not the differentiator. Task structure is.

METR's ongoing time-horizon measurements give the other half of the picture. METR measures the task duration at which a frontier agent succeeds 50% of the time, and finds that horizon growing exponentially — but with two caveats that GTM buyers keep dropping. First, the 50% horizon is a coin flip, not a deployment standard; the 80% horizon is meaningfully shorter. Second, METR's own FAQ is blunt that its tasks are clean, self-contained, and algorithmically scored, and that agent performance drops substantially when scored holistically instead. Outbound is the opposite of a clean, self-contained task. It is messy, human-scored, and the success criterion is whether a stranger chose to reply.

The compounding-error arithmetic

This is not research, it is arithmetic, but it is the arithmetic every long-horizon pitch elides. Take an agent with a given per-step reliability, running a chain of dependent steps where each step's output feeds the next:

Per-step reliability5 dependent steps20 steps50 steps
95%77%36%8%
98%90%67%36%
99%95%82%61%
99.9%99.5%98%95%

A six-week outbound sequence against one account — research, enrich, qualify, draft, send, read the reply, adjust, re-sequence, escalate — is easily fifty dependent steps. At 98% per-step reliability, which would be an excellent result for any current agent on messy real-world work, roughly one in three accounts finishes the run correctly.

This is why the interesting engineering in long-horizon systems is checkpointing, verification, and human approval gates rather than raw capability. It is also why "the agent runs autonomously for weeks" and "the agent is reliable" are close to being in tension. Salesforce clearly knows this — the guardrails and seller-approval language in the announcement exists precisely because unattended fifty-step chains do not survive contact with reality.

What breaks first in a long-horizon outbound run

Mapping the HORIZON failure taxonomy onto outbound gives you a concrete list of what to watch for in a pilot.

1. The goal was wrong on day one

"Rescue my at-risk deals" is a good demo prompt and a bad production objective, because "at risk" is defined by CRM hygiene that is usually two weeks stale. A long-horizon agent takes your objective literally and then spends six weeks executing against it with perfect persistence. Persistence against a bad target is worse than no persistence at all, because it manufactures activity that looks like progress in every dashboard you own.

The failure is not the agent's. It is that the goal was derived from data that only described accounts already in your CRM.

2. The signal decayed underneath the plan

This is the failure mode unique to sales, and the one no general agent benchmark captures. A buying signal has a half-life. Someone who viewed your profile on Monday is a different prospect by Friday. A hiring post for a Head of RevOps is a live signal for maybe two weeks before the role is filled or frozen. A funding announcement is warmest in the first ten days.

A multi-week agent plan built on a Tuesday signal and executed over six weeks will, by week three, be doing careful, well-researched, deeply personalised outreach about something the prospect has stopped caring about. The agent has no way to know this unless the signal layer feeding it is live and continuously re-scored. Durable execution makes this worse, not better: the agent is now durably committed to a stale premise.

3. Memory drift and catastrophic forgetting

HORIZON found memory failures becoming a dominant category as horizons extend. In outbound this looks specific and embarrassing: the agent forgets that the prospect already replied, forgets an objection that was raised in week two, re-sends a variant of a message it already sent, or contradicts a claim a human made on a call. Every one of those is a brand cost you pay in a channel where you get one shot.

4. Nobody owns the output

When an SDR sends 40 messages, a person is accountable for those 40 messages. When an agent runs 200 accounts across six weeks with partial autonomy, accountability diffuses. This is a management problem, not a model problem, and it is the one most likely to sink a pilot quietly.

The rule that makes long-horizon agents work: match horizon to signal half-life

The practical takeaway from all of this is a single design rule. Do not let an agent's autonomous horizon exceed the half-life of the signal the campaign is built on. If the signal decays faster than the plan executes, the plan is wrong before it finishes.

Signal typePractical half-lifeSafe autonomous horizon
Profile view, post engagement2–5 daysHours to a couple of days, then re-check
Competitor complaint or mention1–2 weeksUnder a week per plan cycle
Hiring post for a relevant role2–4 weeks1–2 weeks, re-verify the role is still open
Funding announcement4–8 weeks2–3 weeks per cycle
Job change into a buying role1–3 monthsGenuinely fits a multi-week horizon
Cold ICP-fit list, no signalNo decay, no urgencyLong horizon is possible and mostly pointless

Read the bottom two rows together, because they contain the real insight. The only outbound work that genuinely suits a weeks-long autonomous horizon is work anchored to a slow-decaying signal or to an existing relationship. Everything fast-moving needs short cycles with fresh input. And cold, signal-free lists tolerate any horizon you like precisely because nothing about them is time-sensitive — which is also why they convert badly no matter how long the agent persists.

This is the same conclusion the buyer data keeps pointing at from a different direction. Gartner's research found that 69% of B2B buyers turn to sales reps to validate AI-generated insights, and that sales organisations providing AI-enabled next best actions are 2.6x more likely to achieve commercial growth. AI that decides what to do next from good inputs outperforms. AI operating unattended in front of the buyer is where the confidence gap opens.

How to pilot a long-horizon outbound agent in six weeks

If you are going to run a pilot when Hunter or an equivalent hits GA in November, structure it so the failure modes above surface early and cheaply.

  • Cap the horizon at two weeks before a human checkpoint. Not because the runtime cannot go longer, but because two weeks is roughly where you can still diagnose a wrong plan from its output. Extend only after you have seen clean two-week runs.
  • Pick one signal type, not a mixed list. Run the pilot on a single warm trigger with a known half-life. Mixed-provenance lists make it impossible to attribute failures.
  • Instrument per-step reliability, not just outcomes. Log every agent decision and sample 20 trajectories a week. You want to know where the run diverged, which is precisely what aggregate meeting counts hide.
  • Hard-gate every outbound send for the first three weeks. Approval fatigue is real, so track your approval-override rate. If humans are rejecting more than 10% of drafts in week three, the input layer is wrong, not the copy.
  • Run a human control group. Same signal, same ICP, human-executed. Without it you cannot distinguish agent performance from signal performance, and signal performance is usually the bigger variable.
  • Measure win rate on agent-sourced pipeline, not pipeline volume. Volume is the metric vendors report. Win rate is the metric that tells you whether the pipeline was real.
  • Set a kill criterion before you start. Write down the reply rate, the meeting-to-opportunity rate, and the brand-complaint threshold that end the pilot. Agreed in advance, these are cheap. Agreed in month three, they are political.

The four numbers to instrument

Per-step divergence rate, approval-override rate, signal-to-send latency (how many days pass between the trigger firing and the message landing), and win rate on agent-sourced opportunities versus human-sourced. Signal-to-send latency is the one most teams do not track and the one that best predicts whether a long-horizon run will work — if your median is over five days, your horizon is already too long for most triggers.

Where the input layer fits

Every long-horizon agent argument eventually collapses into the same point: the runtime is downstream of the targeting. Memory, durable execution, and dynamic steering are all machinery for pursuing an objective faithfully. None of them evaluates whether the objective was worth pursuing, and none of them notices that the reason you targeted an account in week one evaporated in week three.

That is the layer worth investing in before November. A signal-based platform like Updately exists to solve the input half of this — watching for profile views, post engagers, competitor mentions, hiring posts, and pain-point discussions on LinkedIn, Reddit and X; scoring them against your ICP; researching the prospect properly; and sending inside safe limits while the trigger is still warm. Feed a long-horizon agent live, scored, decaying-aware signals and the weeks-long horizon becomes an advantage. Feed it a CRM export and the horizon just means it takes six weeks to find out the list was wrong.

The teams that win the next two quarters will not be the ones with the most autonomous agent. They will be the ones whose agents are pointed at the right hundred accounts and re-pointed the moment the evidence changes.

Takeaways

  • Long-horizon AI sales agents are real as of this week. Salesforce's Agentforce long-horizon runtime shipped 11 September 2026 with Hunter as the first outbound agent on it, pilot now, GA November 2026.
  • Reliability collapses abruptly, it does not fade. HORIZON's 3,100-trajectory study found sharp, non-linear performance drops past a horizon threshold, with planning and memory failures becoming dominant, and model differences narrowing to nothing past the breaking point.
  • Compounding error is the constraint. At 98% per-step reliability, a fifty-step dependent chain finishes correctly roughly a third of the time. Checkpoints and approval gates are the fix, not more capability.
  • Match the autonomous horizon to the signal half-life. Profile views and post engagement need same-week action. Job changes and funding tolerate weeks. Never let the plan outlive the premise.
  • Pilot with a two-week cap, one signal type, a human control group, and a written kill criterion. Instrument signal-to-send latency and win rate, not pipeline volume.
  • The runtime is downstream of targeting. An agent that persists brilliantly against the wrong accounts is a more expensive version of the problem you already have.

Ask your vendor for win rate on agent-sourced pipeline and cost per agent-sourced opportunity. Nobody is publishing those yet. The first team that gets them in writing will have a considerably better November than the teams buying on adjectives.