Skip to content
Document Management

79% Say They're Adopting AI Agents. Most Never Reach Production

In PwC's survey, 79% of executives said their companies are adopting AI agents, yet two-thirds said fewer than half their staff actually use them. The gap isn't model quality. It's trust.

ت
تیم فنپینو
August 16, 2026
79% Say They're Adopting AI Agents. Most Never Reach Production

The adoption numbers tell two different stories depending on which stat you read. Gartner predicts that 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from less than 5% in 2025. In PwC’s May 2025 survey of 300 US senior executives, 79% said AI agents were already being adopted in their companies, yet two-thirds said fewer than half of their employees work with agents day to day. Read past the headline and most of that adoption is pilots, demos, and internal tools nobody trusts enough to ship.

That gap isn't about model quality. GPT-class and open-source models both cleared the bar for "good enough" a while ago — on Artificial Analysis data compiled by WhatLLM (June 2026), the best open-weight models trail the closed frontier by only four to five points on the agentic and coding indexes, at a much lower price per token. The gap is about what happens when an agent is wrong in front of an auditor, a regulator, or a customer.

Update, October 2026. Newer data says the same thing with a bigger sample. In Teradata's July 2026 Wakefield Research study of 1,000 senior technology and data leaders across the United States, the United Kingdom, France, Germany, Japan, and Saudi Arabia, only 7% of enterprises had reached the stage where agents execute multi-step workflows with measurable business impact; 68% were still experimenting or developing. The blocker it names is not the model either: 77% of executives said 20% or less of their enterprise data is described and contextualized well enough for an agent to use, and 63% reported no more than a small or emerging return on their agentic AI spend so far.

Why banking beats healthcare on production AI

Production adoption isn't evenly spread. Banking and insurance tend to get agents into production first, while healthcare and government move much more slowly. That's not a coincidence of budget. Banking and insurance already run on audit trails, sign-offs, and traceable decisions; an AI agent slots into a process built for scrutiny. Healthcare and government don't have that scaffolding yet, so an agent that "usually gets it right" is a liability, not a feature.

The EU AI Act's transparency rules for chatbots and AI-generated content have applied since August 2, 2026, and the high-risk system rules follow in December 2027 after the Digital Omnibus delay the Council approved in June 2026. Under Article 99, fines reach €35 million or 7% of global turnover for banned practices and €15 million or 3% for most other breaches. Every organization now has a concrete reason to ask an uncomfortable question before deploying an agent: if this answer is wrong, can we show exactly why it said what it said?

The fix isn't a smarter model — it's a narrower one

The deployments that made it past the pilot stage share a pattern: they're task-specific, not general-purpose. A reporting agent that builds one kind of report can be audited end to end, so it ships. A clinical documentation assistant gets used for the same reason: the task is bounded and it leaves a paper trail. The general assistant that does a bit of everything rarely gets that far.

That's the same principle we built FanMind around, before "agentic AI" was the headline of the year. FanMind doesn't answer from memory or guesswork — every response is grounded in an organization's own uploaded documents, with a citation to the exact document and page. If the answer isn't in the documents, it says so, and logs the question so a content admin can close the gap instead of the model quietly inventing one. Access control follows the organization's existing permissions, not a flat "anyone can ask anything" model. It's less flashy than a general-purpose agent that promises to do everything — but it's the version that can actually go into production, because every answer is traceable back to a source a human can check.

If your AI initiative is stuck in the 68-point gap between "we adopted it" and "we trust it enough to run it," the question worth asking isn't which model to upgrade to. It's whether the agent can show its work.

On the coding side, we wrote up how we run Research, Plan, Implement with Claude Code subagents. Agents also need their tool servers checked, and our audit of the MCP servers we run after the Deadbugz campaign found three of them running without a version pin.

Model size is not the fix for this gap either; we looked at why in how a 7M-parameter model outscored much larger LLMs on ARC-AGI.

See FanMind

A grounded, citation-based AI assistant built to actually reach production — not stay a pilot.

View product

Share This Article