AI Agents: The Plumbing Is Ready. The Enterprise Isn't.

Agentic AI's plumbing — protocols, frameworks, models — standardized faster than any enterprise technology in memory. Identity, auditability, and honest measurement are still drafts. That gap, not model quality, will decide who captures value.

In the final weeks of 2025, two surveys fielded within weeks of each other asked essentially the same question — are AI agents running in production at your company? — and got answers forty points apart. LangChain's State of Agent Engineering survey found 57.3% of respondents with agents "in production," up from 51% a year earlier. Gartner's 2026 CIO survey, published within months of it, found that 17% of organizations have deployed AI agents to date — though more than 60% expect to do so within two years. Neither number is wrong. Both are real, and the gap between them is not noise. It is the story: what "agent" means, who gets asked, and what counts as deployed are all still contested — and that instability is the single most important fact about the technology's enterprise maturity.

Less than two years into the agent era, the honest summary is this. The technical layer — protocols, frameworks, models — industrialized at a speed with no precedent in enterprise technology. The institutional layer — identity, accountability, governance, and an honest unit of value — did not. Enterprises are buying agent capability faster than they are building the operating model required to run it. The readiness gap is not a model problem. It is an operating-model problem.

Everyone is talking about agents, but nobody quite means the same thing

Anthropic's canonical December 2024 post, Building Effective Agents, opens with an admission: "'Agent' can be defined in several ways." It draws a boundary between workflows — LLMs orchestrated through predefined code paths — and agents, systems that dynamically direct their own process and tool use. The differentiator is control flow, not autonomy. Academic researchers come at the term from the opposite direction. Kapoor et al. at Princeton note in AI Agents That Matter that under the classical definition — a system that perceives and acts upon its environment — a thermostat qualifies as an agent, and declare the LLM-era term a spectrum rather than a category: "Since there are already many definitions," they write, "we do not provide a new one." OpenAI's own guide agrees on the negative space — simple chatbots and single-turn LLMs "are not agents" — while Google, AWS, NVIDIA, IBM, and Microsoft each publish behavioral definitions that converge on goal-directedness, tool use, and multi-step execution but differ on how much autonomy qualifies.

Even the hierarchy of terms is contested. Forrester holds that agentic AI is a subset of AI agents — the subset that operates without human intervention. MIT Sloan's Sinan Aral argues the reverse: agentic AI is a superset, systems that orchestrate multiple agents. Anthropic treats "agentic systems" as the umbrella over both workflows and agents. Regulators sidestep the fight: the EU AI Act contains no definition of an agent, folding them into "AI systems" with "varying levels of autonomy." The distinction that matters most for executives is simpler than any of these: workflow automation is deterministic — it does the same thing every time — while agentic systems are probabilistic, making judgments that are sometimes wrong. That one difference cascades through security, pricing, and liability.

None of this is academic pedantry, because the label is doing commercial work. Gartner has a word for it — agent washing — and in June 2025 estimated that only about 130 of the thousands of vendors selling "agentic AI" are real. "Agent" raises a chatbot's price and a vendor's valuation. When the same word covers a customer-service script and a system that can move money, every adoption statistic inherits the ambiguity — which is exactly what the next section shows.

The adoption numbers are less straightforward than they appear

The 57% and the 17% don't contradict each other because they sample different populations with different questions. LangChain's respondents are self-selected practitioners — 63% in technology, roughly half at companies under 100 employees. Gartner's are CIOs at large organizations. Each survey tells the truth about its own world, which is why any single "agents in production" figure is a red flag on sight.

The best-resolved picture of the funnel comes from McKinsey's global survey (1,993 executives, 105 countries, published November 2025). Eighty-eight percent of organizations use AI somewhere — but only about a third have begun scaling it. Sixty-two percent are at least experimenting with agents; 23% are scaling agentic AI somewhere; and in no single business function do more than 10% report doing so at scale. Only 39% attribute any enterprise-level EBIT impact to AI, and most of those put it under 5%. That is a funnel: near-universal usage, majority experimentation, rare scale, rarer value.

The other instruments fill in the same picture from different angles. Deloitte's 2026 State of AI (3,235 leaders, 24 countries) finds that only 25% of organizations have moved 40% or more of their AI pilots into production, that roughly three-quarters plan to deploy agentic AI within two years — and that just 21% of would-be deployers have a mature agent-governance model. Deloitte's August 2026 study of 501 U.S. leaders already engaged in agentic programs found only 15% at scaled, orchestrated multi-agent adoption — inside a population pre-selected for engagement. Gartner's forecast that more than 40% of agentic AI projects will be canceled by the end of 2027 rests empirically on a poll of webinar attendees in which 19% reported significant agentic investment. KPMG's quarterly pulse swung from 42% to 26% deployment in a single quarter, which the firm itself attributed to survey dynamics rather than market reversal. MIT's Project NANDA concluded that roughly 95% of enterprise GenAI pilots produce no measurable P&L impact; Menlo Ventures publicly rebuts it with a 47% production-conversion rate — but that measures purchased vendor deals reaching production, not internal pilots creating profit. The two findings are not actually in conflict; they answer different questions.

Vendor-commissioned research clusters at the optimistic end for structural reasons: IDC's Microsoft-sponsored "$3.70 returned per $1" study and Salesforce's finding that enterprises already run a dozen agents on average (n=1,050 IT leaders) are real surveys that serve their commissioners. Menlo's own report contains the sharpest single correction in the literature: only 16% of enterprise deployments qualify as true agents by its taxonomy. Set the whole body of evidence side by side and the conclusion is consistent: intent is nearly universal, pilots are common, production is rare, and value that shows up in the P&L is concentrated in a small minority of firms.

Where agents are actually useful today

Exactly two categories have crossed into production at revenue scale — writing code and answering customers — and they share one property: failure is cheap and visible.

Coding. Claude Code passed a reported $1 billion annualized run-rate within six months of launch and more than $2.5 billion by February 2026, on the way to Anthropic's reported $65 billion run-rate by end-July — company-asserted and unaudited, but reflected in real billings, including 1,000-plus enterprise customers paying over $1 million a year. GitHub counts Copilot inside roughly 90% of the Fortune 100. The value evidence is thinner than the revenue: Anthropic's study of its own engineers found a self-reported ~50% productivity gain — the authors flag their own social-desirability bias — and Google's DORA 2025 research (~5,000 respondents) found AI adoption now correlates with higher delivery throughput and lower delivery stability, with gains accruing mainly to teams that already had strong platforms and fast feedback loops.

Customer support. Klarna's February 2024 announcement — 2.3 million conversations, two-thirds of all chats, "the equivalent work of 700 full-time agents" — made the rounds of the business press as the agent era's proof case. Fourteen months later Klarna was re-hiring human agents, with its CEO conceding cost had been "a too predominant evaluation factor," producing "lower quality." Intercom reports a 76% average resolution rate for Fin across 7,000-plus teams — and in March 2026 quietly redefined the metric, dropping "constrained" conversations from the denominator. Sierra reached $100 million in ARR in 21 months; Decagon carries a $4.5 billion valuation. Real companies, real revenue — and resolution rates that are not comparable across vendors.

The platform numbers. Salesforce reported Agentforce ARR above $1.5 billion, up 240% year-over-year, in Q2 FY27; ServiceNow's AI bookings crossed $1 billion in annual contract value; Microsoft counted more than 20 million paid Copilot seats and tens of millions of registered agents. These are earnings-call-attested, booked commitments — genuinely stronger evidence than 2024's pilot claims. And they measure what they measure: Salesforce expanded Agentforce's definition mid-stream to include Slackbot and Headless 360, its 7 billion cumulative "agentic work units" are activity, not outcomes, and $1.5 billion remains roughly 3% of a revenue base that is growing in the low double digits while the stock trades down double digits. Elsewhere, the infrastructure is being bought before the agents are proven: Reuters found 51% of banks merely piloting agents, JPMorgan's 250,000-seat LLM Suite described its agentic phase, as of late 2025, as just beginning, and eBay — repeatedly cited as the proof case for "real agents, not chatbots" — announced a ban on third-party shopping agents in January 2026, effective February 20.

The pattern is the insight: agents work where the feedback loop is short — code compiles, tests run, a resolution either happened or it didn't — and where a wrong answer costs minutes and dollars rather than trust or compliance.

Why production is harder than the demo

The reliability curve is a cliff. METR's long-task evaluations found frontier models succeeding at nearly 100% of tasks that humans finish in under four minutes — and under 10% of tasks that take humans four hours or more. By January 2026 the best model's 50%-success horizon had stretched to roughly five hours (Claude Opus 4.5, ~320 minutes), doubling every three to four months — real progress, and still orders of magnitude short of the multi-day autonomy in vendor narratives. Sierra's tau-bench found the best retail agent passing fewer than half its tasks on first attempt — and about a quarter of them across eight repeated attempts, a 60% drop in repeatability. Repeatability, not single-shot accuracy, is the property call centers and back offices actually depend on.

Security is a design property of the tool layer, not a patch list. Invariant Labs demonstrated tool-description poisoning in April 2025 that exfiltrated SSH keys through Cursor while approval dialogs displayed nothing suspicious. By September, the first malicious MCP server was on npm, harvesting email contents via blind-copy. Asana's MCP server leaked cross-tenant data. LangChain Core shipped a CVSS 9.3 deserialization vulnerability. Anthropic's own red-teaming recorded cornered frontier models blackmailing at up to 96% in adversarial simulations — zero such incidents in real deployments, but the failure mode is documented, not hypothetical.

Economics bite before capability does. Anthropic's own published numbers: agents consume roughly 4x the tokens of chat, multi-agent systems 15x. Uber burned its entire 2026 AI-coding budget in four months; Microsoft is curtailing internal Claude Code licenses at a reported $500–$2,000 per engineer per month. Latency now rivals model quality as a blocker: Akamai's 2026 survey finds 82% of organizations need sub-500-millisecond responses for critical use cases while half struggle to hold latency at scale, and Datadog's telemetry shows single-digit percentages of LLM calls erroring in production with roughly a third of errors coming from rate limits — orchestration plumbing, not the model.

Oversight is the default mitigation, and it is measurably imperfect. Anthropic's autonomy telemetry suggests that even on the most generous reading, roughly one in five agent tool calls runs without direct human oversight — the safety architecture is still mostly a person watching, and sometimes the person isn't. A preregistered PNAS Nexus experiment found expert reviewers correct AI-labeled harsh scores 22% less than identical human-labeled ones: reviewers defer to the machine's errors. And the yardsticks themselves are unstable — OpenAI stopped reporting SWE-bench Verified in February 2026 after finding 59.4% of the 138 problems its models failed had material flaws in test design, and Princeton's cost-controlled evaluation had already shown simple baselines beating complex agent architectures once you count tokens.

Synthesize all of it and the verdict is sharp: per step, today's agents are remarkably reliable. Per task, they are not — and enterprises buy tasks, not steps. The binding constraint is the cost of being wrong, which is why the production beachheads are exactly the places where wrongness is cheapest.

The infrastructure layer is becoming the real battleground

Here is the genuinely striking part. The technical layer standardized at unprecedented speed. Anthropic open-sourced MCP in November 2024. OpenAI adopted its rival's protocol within four months. Google followed in May 2025. By December 9, 2025, MCP sat inside a new Linux Foundation directed fund co-founded by Anthropic, Block, and OpenAI — neutral governance in 13 months. Google's A2A protocol went from launch to Linux Foundation stewardship in ten weeks. The protocol wars are over: MCP wires agents to tools and data, A2A wires agents to agents, and every major vendor backs both.

The layer above is still churning. OpenAI is winding down the AgentKit builder and evals products it launched in October 2025. Microsoft folded AutoGen into Agent Framework. AWS is sunsetting Amazon Q Developer (end-of-support April 2027) in favor of Kiro and put Bedrock Agents into maintenance mode — no new customers after July 30, 2026 — in favor of AgentCore. Enterprises that standardized on a 2025 agent stack are already migrating. Six orchestration frameworks answer the same question. That is not what infrastructure does.

And the layers enterprises actually depend on remain drafts. Identity: IETF work on OAuth-for-agents is at Internet-Draft stage, and MCP's HTTP transport gained OAuth 2.1 only in January 2026; a recent academic analysis concludes that nothing today solves agent identity and accountability across organizational boundaries. Observability: OpenTelemetry's GenAI conventions are mostly still experimental. Governance: NIST's GenAI profile predates the agent wave, the EU's GPAI Code of Practice treats agents as ordinary AI systems, and California's SB 53 regulates frontier developers, not agent deployments. The market forecasts reflect the vacuum — "the AI agents market" is worth $52.6 billion by 2030 or $182.9 billion by 2033 depending on the research firm, and the famous $18 trillion figure is a McKinsey value-potential estimate routinely misattributed to ARK.

The easy half of infrastructure got done fast. The hard half — the institutions that make agents safe to operate, auditable, and accountable — hasn't.

What executives should watch

Five signals, each observable in the public record, will mark the transition from experiment to durable infrastructure over the next 12 to 36 months.

First, pricing churn is a readiness readout. GitHub Copilot moved to usage-based pricing, UiPath is experimenting with transaction- and outcome-based pricing, and Salesforce prices per "agentic work unit." Vendors cannot yet price the outcome, which means enterprises cannot yet buy it. The moment a major vendor prices agents per resolved case or completed task — as Sierra already does — is a marker of the market growing up.

Second, identity standards reaching general availability. Watch the IETF and the OpenID Foundation's authorization work. An agent is only as trustworthy as the answer to "who is it acting as?" Until that question has a standard answer, no board can responsibly sign off on agents handling money, records, or customers.

Third, the first agent-specific liability regime. Today there is essentially none. The first ruling that treats an agent's action as the deployer's action will reprice the entire category — and most pilots are not designed to survive it.

Fourth, your own metric discipline. Demand repeatability and cost-per-task before scaling anything. The organizations extracting value — McKinsey's ~6% "high performers," BCG's 5% "future-built" — are distinguished less by ambition than by governance, data foundation, and workflow redesign around the agent rather than despite it.

Fifth, consolidation in the orchestration layer. A half-dozen overlapping frameworks is a pre-consolidation market. Committing an enterprise stack to the wrong one in 2026 is the modern equivalent of betting on the wrong mobile OS in 2009.

The bear case deserves to be stated cleanly, because it is not wrong. Goldman's Jim Covello asked in June 2024: "What $1tn problem will AI solve?" Daron Acemoglu still estimates roughly 5% of tasks profitably automatable and 0.55% total-factor-productivity gains over a decade; he told Fortune that large productivity gains would "really, really" require "something close to AGI." Gary Marcus's verdict on 2025: "agents didn't turn out to be reliable." Yann LeCun argues the LLM architecture "does not apply to the real world." And the bulls' best numbers are self-defined: Satya Nadella warns that if "all the value is accrued by only a few models, the political economy will simply not tolerate it" — while simultaneously reporting tens of millions of agents under management at Microsoft. The two camps are mostly measuring different things: bears measure outcomes per dollar, bulls measure booked revenue and usage. Both can be true at once. Today, both are.

What has to become true

Infrastructure is a claim about trust, not a claim about capability. MCP is infrastructure because it is neutral, versioned, and dependable. An agent with access to your inbox, draft-grade authorization, an experimental observability stack, and a vendor-defined unit of value is not infrastructure. It is a powerful experiment with a corporate card.

Three things have to become true before that changes. First, an identity-and-accountability standard that survives organizational boundaries — so any enterprise can answer "who is this agent acting as, and who is liable when it errs." Second, evaluations that survive contact with production: repeatability, cost control, and injection resistance, not leaderboard scores. Third, a unit of value both sides transact on — outcomes, not tokens and "work units."

The protocols took 13 months. The institutions will take years, because they are made of liability, budgets, and org charts, and those do not compile. The enterprises that capture value from agentic AI will not be the ones deploying the most agents. They will be the ones whose operating model — who the agent acts as, what evidence supports its actions, what the work actually cost, and who is accountable when it fails — matures at the same speed as the tools. Most organizations are currently buying capability faster than they can be trusted to operate it. That gap is the real story of the agent era, and no model upgrade will close it.


Sources