45% of the S&P 500 are running AI pilots. Eleven percent are real. Here's the gap.
Almost every enterprise AI number you read comes from a survey someone paid to like AI. In August, MIT researchers stopped asking companies how AI-forward they feel and read what they told the SEC under penalty of perjury. The picture is humbler — and the three independent datasets that landed in September agree with it.
The 10-K trick
A team from the MIT Initiative on the Digital Economy and MIT FutureTech ("AI Adoption in S&P 500 Firms", Yu, Fleming, Hampton, Combemale & Thompson) classified AI-related language in the annual 10-K filings of 510 companies across ten years. The methodology's whole point: 10-Ks are legally binding documents — companies cannot make materially false statements in them — so they filter the marketing out.
- 11% of S&P 500 firms had AI deeply integrated into core business processes by end of 2025. (A further 10% use AI in production of goods or delivery of services — 21% at "meaningful" levels in total.)
- 45% are running pilots — and the researchers themselves note many of these pilots will never reach production.
- Two-thirds of deep integration is at technology companies. In financial services: 68% pilot heavily, about 4% deeply integrated. For banks specifically: ~85% have pilots, essentially none have reached deep integration.
- The money trail is a J-curve: early adopters run 2–3 points lower margins (infrastructure, retraining, workflow disruption); firms that reach deep integration in goods/services report ~5% margin improvement.
The banks number deserves a second look, because that is the most regulated, most document-heavy cohort in the economy: nearly every major financial institution is experimenting with AI, and almost none has let it into the production line of how it actually operates. The blocker in the filings' language is rarely model quality. It is what the company is allowed to let the model touch.
Two more surveys say the same thing
UiPath's Wakefield study of 590 C-suite and IT leaders (published around FUSION 2026, September): two-thirds of enterprises have agentic AI embedded or adopted among select teams — but only 29% say orchestration is fully embedded in workflows. And the leaders cited as what's blocking scale line up with the pilots-to-production gap almost point for point: data quality and readiness (38%), integration with existing systems and workflows (37%), governance and compliance (33%).
KPMG's Q3 Global AI Pulse (2,131 senior leaders, 20 countries, September 24) adds the money: average planned AI investment rose from $186M to $210M per organisation over the next 12 months — while only 12% consistently assess AI value against cost. Among the minority reporting established ROI, 86% have what KPMG calls a formal AI harness layer — the controls sitting between models and business use — versus 31% at the experimentation stage.
Three surveys that never coordinated with each other, three continents of respondents, one story: spending is exploding at the pilot layer and stalling at the governance layer.
What the gap actually is
Read together, the datasets put a price on the distance between "we tried it" and "it runs the business." It is not a model problem — 45% piloting proves capability is cheaply available. It is four unglamorous problems:
- Data can't be touched: customer names, deal terms, client files, health numbers. The first gate every regulated firm hits before a model sees anything.
- Nobody can replay the decision: if a regulator asks "what did your AI see and say in March," "we think it was fine" is not an answer. The KPMG cohort that made ROI work treats accountability as a C-suite object (53% put it at C-level or above).
- The tools don't talk to each other (37% of UiPath's blockers), and
- Nobody measures the cost side — 12% consistent value-vs-cost assessment against $210M average planned budgets is the loudest number in all three studies.
The J-curve is where those costs land: the bottom of the J is exactly the stretch where governance, integration and measurement are paid for but benefits haven't compounded. Firms that clear it show up as margins, not headlines.
What it means if you're at a 20–200 person firm
The MIT study only counts the S&P 500, but every vendor in the other two studies says the same blockers hit mid-market first and hardest: no platform team to build the harness, no compliance budget to build the audit trail, no data engineers to wire the integrations. That is precisely the gap between "every lawyer/engineer/accountant with an AI account, at personal-chatbot risk level" and "deep integration." The 11% that made it did not start by integrating everything. They started by letting one trusted path — inputs masked, outputs logged, spend metered — prove itself in production.