Deployment stories · September 25, 2026 · 7 min read

"Almost 90% of our agent prototypes never shipped." Amazon said it out loud.

Swami Sivasubramanian, AWS vice president of agentic AI, stood on stage at HumanX in Amsterdam this week and volunteered the statistic every enterprise quietly already knows about its own AI graveyard. Then he walked through the autopsy — five root causes, found by engineers embedded at customers over six months — and the rebuild that followed. It is the most useful enterprise-AI story of the quarter precisely because the confession came from inside the company with the most to sell you.

The confession and the base rates

The numbers, as TNW reported from the talk (September 24): ~90% of the agent prototypes Amazon teams built two years ago never reached production. External analysts he cited put it only slightly kinder — 17% of organizations have successfully deployed AI agents, and only 7% can measure the return. OpenAI's Colin Jarvis, at the same conference, put the consensus bluntly: enterprise AI is stuck on deployment, not models.

The five causes of death

AWS's forward-deployed engineer program — actual engineers sitting inside customer projects — spent six months tracing failed agent projects back to five patterns. They are worth reading as a diagnostic checklist for any firm with AI pilots in flight:

Notice what is not on that list: model capability. Two of the five causes are measurement, two are governance/organizational, one is problem choice. The AI Pilot Papers from MIT (where 45% of the S&P 500 are piloting and 11% have deep integration) and KPMG's Q3 pulse (86% of organizations with real ROI operate a formal harness layer) land on the same verdict from different directions.

What Amazon did about its own graveyard

Three moves from the talk are worth knowing even if you never touch AWS:

1. Consolidate onto one approved path. Amazon teams had built agents on EC2, EKS, SageMaker and other services — every stack a new surface for security review. Amazon collapsed the paths: Bedrock for inference, AgentCore for hosting, "a single path approved by security." The internal complaint that drove this — security and infrastructure teams strained by a hundred bespoke agent stacks — is the real reason most enterprise AI reviews feel slow.

2. Let a thousand flowers bloom, then measure them. More than 100,000 Amazon engineers now use Kiro, the agentic coding tool; measurement systems flagged which agent projects were dead weight, and the survivors got merged — always-on agents, persistent memory, multi-agent coordination prototypes fused into Kiro Crew, which 39,000 employees adopted within 30 days of internal launch.

3. Fence the agent, widen the fence as it proves itself. The Strands framework feature he demoed is the cleanest metaphor of the quarter: a deterministic layer outside the agent that governs which tool calls it may make, expanding as the agent earns trust. Governance as a ratchet, not a gate.

The counter-example that proves the model works: Toyota

While AWS was describing the failure rate, Toyota Motor North America published its success ledger (August 24) — a ~35-person enterprise AI team running 50+ agents in production, and the before/after is the whole thesis: a new agent used to take 6 months and 6 engineers; now it takes 4 days and 1 engineer. GearPal turns a 5–6 hour machine diagnosis into 2–3 minutes on the factory floor. R&D GPT compressed research reviews from ~3 years to ~1. Each manufacturing use case is projected at six figures of annual savings, and the team's platform lead calls their observability stack "the Andon board for AI applications" — Toyota's factory principle made digital: everyone can see what's running, what's failing, and what it costs.

The difference between Toyota and the dead 90% is not better models. It is that Toyota built the boring machinery first — one platform path, permission-gated data access (if you lack SharePoint access, the agent stays blind too), every agent traceable, and ROI per line-item — and then scaled agents aggressively. Which is exactly the ratchet: strong fences, measured wins, widened gates.

The 200-person version of this story

Amazon and Toyota have platform teams, internal marketing for pilots, and budget for J-curve dips. A 200-person law, accounting or fintech firm has none of that — but it faces the identical math, because the five causes of death don't scale down; they get worse. The ratchet for a smaller firm is a sequence of cheap, boring steps:

The 90% figure is not a warning that agents fail. It is a map of where they fail — and every item on the map sits before the model, in the layer around it.

One governed path is the whole head start.

Start free — $5 credits, no card