(1/3) 95% of GenAI Pilots Show No Return. The Models Were Fine.
The most quoted AI number of the past year is a failure rate: 95% of enterprise GenAI pilots deliver no measurable return. I've watched that statistic do a full tour — board decks, LinkedIn doom threads, budget-freeze memos — and almost nobody quoting it has read what it measures. Because read properly, the number acquits the technology. The pilots aren't dying in the model. They're dying in deployment, inside environments the demo never met.
I've spent two decades on the operations side of software — enterprise CI/CD, incident response, the unglamorous plumbing that decides whether something ships or rots in staging. I have sat in the meeting where a pilot goes to die, and the cause of death is never "the model wasn't smart enough." It's an access review that took eleven weeks. It's a workflow owner who was never asked to change the workflow. This post is about what the 95% actually says, and why the market's answer to it is a hiring pattern, not a better model.
What MIT measured — and what the headline dropped
The figure comes from The GenAI Divide: State of AI in Business 2025, published by MIT's Project NANDA in July 2025: a review of 300+ publicly disclosed AI initiatives, 52 structured interviews, and 153 survey responses from senior leaders, against an estimated $30–40 billion of enterprise GenAI spend. The finding: 95% of the pilots studied produced no measurable P&L impact in the study window. The other 5% were extracting millions.
Two things the headline version omits. First, "no measurable P&L impact yet" is not "the technology failed" — plenty of honest enterprise IT categories would post ugly numbers on a six-month P&L test. Second, and more damning for how the stat gets used: the report's own diagnosis is not model capability. The authors point at the absence of learning, integration and contextual adaptation — tools that never absorbed the organization's context, workflows that never absorbed the tool. The report doesn't say AI failed. It says deployment did. If you quote the 95% to argue the models aren't ready, you're citing a study that argues the opposite.
The pilot dies crossing into your environment
The demo-to-production gap isn't a metaphor; it's an itemizable list. The pilot ran on a curated dataset with an admin token and a friendly user. Production means:
- Identity and access. SSO, least-privilege service accounts, an IAM review with a queue. The agent that "just needs read access" needs it to four systems with three different owners.
- Data that lives where data lives. Not the clean export — the CRM with a decade of schema drift, the warehouse where the golden table is golden except on Mondays.
- APIs that were never designed to be driven by software. No idempotency, no sane rate limits, error messages written for a human who can shrug.
- A failure mode someone has to own. The on-call rota has to absorb a new class of incident before the first one happens, not after.
- A workflow with a human in it who has to work differently from now on — and who was told about the pilot in the same email that announced it was live.
None of those items is model capability. Every one of them is engineering and organizational work, and every one is invisible in the demo. The 5% who crossed MIT's divide didn't have better models; the report's success stories share deep integration and adapted workflows — they did the list.
We've seen this movie: DevOps was a tooling purchase once, too
Around 2015 I watched enterprises "adopt DevOps" by buying the toolchain. Jenkins licenses, an artifact repository, a dashboard the CFO liked. Eighteen months later the release cadence hadn't moved, and the retrospective blamed the tools. The companies that got the outcome had done something structurally different: they put engineers inside the delivery workflow with a mandate to change the process, not just install software next to it.
Swap the nouns and it's this year's story. A GenAI pilot bought as a tool and installed next to the workflow produces a demo. The same capability, deployed by someone embedded in the workflow with permission to rewire it, produces a P&L line. The technology was roughly equal in both cases a decade ago, and it's roughly equal now.
The market already priced this in: FDE hiring grew 1,165% in a year
While the 95% was making the rounds as evidence of AI's failure, the companies selling AI were reading it correctly and hiring against it. Forward-deployed engineer roles grew 1,165% year-over-year heading into 2026, per Live Data Technologies' placement data — Perspective AI's analysis of 1,000 FDE job posts tracks the same surge, with postings up roughly 800% in a single nine-month stretch of 2025. Palantir coined the role; OpenAI and Anthropic made it their enterprise-revenue motion; and 59% of the companies hiring FDEs today are Seed through Series A — the surge isn't a big-lab luxury, it's the application layer copying what worked.
A disclosure before I lean on this: deploying engineers this way is my business, so discount my enthusiasm accordingly — then check the hiring data, which isn't mine. An FDE is the deployment list from two sections up, given a job title: an engineer stationed inside the customer's environment, backed by an organization, scoped to a mission with a defined end. I've written before about why the doctrine matters more than the title — forward, backed, mission-scoped. The relevant point here is simpler: when the bottleneck moved from building models to deploying them, the labs didn't respond with a whitepaper. They responded with headcount, at four-digit growth rates. Hiring patterns are the most honest signal a market produces, because they cost money.
What the 95% doesn't license you to conclude
Fairness cuts both ways, so two concessions. If your use case is genuinely contained — clean data, one system, a workflow you own end to end — your own team integrating an API is the right move, and hiring deployment specialists for it would be theater. And the 95% itself deserves skepticism in the other direction: a six-month window for measurable P&L impact is a bar that mature technologies routinely miss. Some of those "failed" pilots will quietly compound into next year's wins. The statistic is a snapshot of a divide, not a verdict on a technology — on either side of the argument.
What it does license: if your pilot has been "two weeks from production" for two quarters, the missing ingredient is almost certainly not a better model, and waiting for the next release won't fix an IAM queue.
What I'd do this quarter with one pilot budget
- Pick one workflow with a P&L line attached. Not a platform, not an "AI enablement layer" — one process where success shows up in a number someone already reports.
- Put an engineer inside the workflow, yours or embedded from outside, with an explicit mandate to change the process — not just to wire the API.
- Give the pilot production constraints from week one. Real IAM, real data, a named owner on the on-call rota. A pilot that runs with an admin token isn't a pilot; it's a rehearsal for a play that will never open.
- Write the exit into the kickoff doc. The adoption metric that means "scale it" and the date you'll read it. Deployment without a defined end is how pilots become furniture.
- Kill it on the metric, not on the vibes. The 95% is full of pilots that were never measured, which means they were never really deployed.
The rest of this series gets concrete: what a forward-deployed engineer's week actually contains, and the economics of the model against classic staff augmentation.
The 95% isn't a verdict on AI. It's the invoice for treating deployment as an afterthought — and the market is already paying it, one forward-deployed hire at a time. If the workflow you need shipped is waiting on that kind of engineer, that's the role we deploy.


