HANDOFF
Reporting on autonomous software

78% of enterprises run an agent pilot. 14% ever scale one.

Surveys put agent pilot failure at 86 to 89 percent, and the most-cited blocker is not the model. It is evaluation infrastructure nobody budgeted for.

By , Evaluation Correspondent Published 8 min read

Key takeaways

  • A March 2026 survey of 650 enterprise technology leaders reported that 78% run at least one agent pilot while only 14% have scaled one to organisation-wide use.
  • Three separate 2026 studies put the pilot failure rate between 86% and 89%, which is a narrow enough band to take seriously despite different methodologies.
  • The most-cited blocker is evaluation gaps at 64% of leaders, ahead of governance friction at 57% and model reliability at 51%. The model is the third problem, not the first.
  • Sector spread is wide: financial services reported 21% production deployment against healthcare at 8%.
  • A pilot has one thing production does not: someone watching it. Evaluation infrastructure is what replaces that person, and it is the line item that never makes the pilot budget.
Advertisement

Does ChatGPT recommend your competitor instead of you?

Check how AI assistants describe your company, and whose name they give when someone asks for a recommendation in your category. Run an audit today, from $19.

Audit your AI visibility at EntityRise.ai →

Seventy-eight percent of enterprises are running an agent pilot. Fourteen percent have scaled one to organisation-wide use. That is from a March 2026 survey of 650 technology leaders, and three other 2026 studies land in the same place, putting pilot failure between 86 and 89 percent.

What makes the number interesting is not its size. It is what the same respondents said was stopping them.

What is actually blocking these pilots?

Not the model. The model comes third.

Blockers to production deployment, as reported by 2026 enterprise surveys. Percentages are the share of leaders naming each blocker, so they do not sum to 100. Not our own research.
BlockerShare naming itWho owns fixing it
Evaluation gaps64%Platform team, usually unstaffed for it
Governance friction57%Risk and legal, usually consulted late
Model reliability51%The vendor, mostly outside your control

The ordering matters more than the values. Teams are not stuck because the models cannot do the work. They are stuck because they cannot demonstrate that the work was done correctly, at a standard anyone will sign off on.

That is an evaluation problem wearing the costume of a capability problem.

Why does a pilot succeed where production fails?

Because a pilot has a person watching it.

During a pilot, someone reads the outputs. They notice when an answer looks wrong, they re-run the odd case, they quietly correct the ones that embarrass them. None of this is recorded as a control, and all of it is doing the work that evaluation infrastructure is supposed to do later.

Then the pilot is judged a success and the same system is pointed at ten times the volume, where nobody is reading anything. The failure rate did not change. The observer left.

The line item nobody budgets

A pilot budget covers model access, some engineering time and a demo. It does not cover a held-out task set, a grader, a trace store or an owner. Those are the four things that turn a demo into a service, and each of them is weeks of unglamorous work with no visible output.

Which is why 64% of leaders name evaluation and 14% ship.

How wide is the spread between industries?

Wide enough to suggest the constraint is organisational rather than technical.

Financial services reported 21% production deployment. Healthcare reported 8%. Both sectors have access to identical models. What differs is that financial services had already built document processing and compliance automation, which means it already had the boring apparatus: labelled outcomes, audit trails, people whose job is to check.

Healthcare’s gap is usually explained by regulation, and that is part of it. The other part is that clinical workflows rarely produce a clean pass or fail signal you can grade automatically.

What would move a pilot across the line?

Four things, in the order they pay off.

What to build before a pilot is judged. This is our analysis of the reported blockers, not measured data.
Build thisReplacesTypical effort
Held-out set of 100 real tasks with known outcomesThe person reading outputs during the pilot1 to 2 weeks
Automated grader running on every changeSpot checks nobody records1 week
Trace store queryable after the factReconstructing failures from memoryDays, if you already run tracing
A named owner with pager dutyShared responsibility, which is noneA decision, not a build

The last row costs nothing and is skipped most often. An agent without an owner is a system whose failures are everybody’s problem, which in practice means they are logged and not read.

The question that predicts whether your pilot ships

Not “did it work in the demo”. Ask: if this system started failing 20% of the time tomorrow, how long before we noticed, and who would be told? If the answer involves a customer complaint, the pilot is not ready regardless of how the demo went.

Frequently asked questions

What percentage of AI agent pilots reach production?

Reported figures cluster between 11% and 14%. A March 2026 survey of 650 enterprise technology leaders put organisation-wide scaling at 14%, while a separate 2026 report put production at scale at 12% against 97% of executives having deployed something.

Why do enterprise AI agent pilots fail?

Surveys put evaluation gaps first at 64% of leaders, governance friction second at 57% and model reliability third at 51%. Analyses of the same data consistently describe the gap as organisational rather than technical: no evaluation infrastructure, no monitoring, no named owner.

Which industries deploy AI agents most successfully?

Financial services reported the highest production rate at 21%, attributed to earlier investment in document processing and compliance automation. Healthcare reported the lowest at 8%, which the same reporting attributes to regulatory complexity and risk aversion in clinical workflows.

What does evaluation infrastructure for an agent actually consist of?

A held-out set of real tasks with known-good outcomes, an automated grader that runs on every change, a trace store you can query after the fact, and an alert when the pass rate moves. None of it is exotic. All of it takes weeks that pilots do not schedule.

Method and sources

  1. March 2026 survey of 650 enterprise technology leaders, reported in 2026 industry analyses. Source for 78% pilot adoption and 14% organisation-wide scaling.
  2. Three 2026 studies attributed to McKinsey, Gartner and a cross-sector analysis by the AI Governance Institute, reported as placing pilot failure between 86% and 89%.
  3. Composio 2026 agent report, cited for 97% of executives deploying agents against 12% reaching production at scale.
  4. Blocker rankings of evaluation gaps 64%, governance friction 57% and model reliability 51%, and sector figures of 21% for financial services against 8% for healthcare, as reported in 2026 analyses of the above surveys.
  5. We have not seen the underlying survey instruments or sample frames for these studies. Figures are cited as reported and should be read as directional rather than precise.
PN

, Evaluation Correspondent at Handoff

Covers evaluation for Handoff: how teams measure whether an agent works, and how those measurements go stale.