Stop Counting Pilots, Start Counting Deployments

A firm spends $300 million annually across seventy new product development projects. Five reach commercialization. Those five deliver five percent of total revenue. The Production and Operations Management researchers who documented this case did not frame it as bad luck. They framed it as a portfolio design failure — too much selection from pre-existing candidate lists, too little alignment between what gets funded and what the business actually needs to accomplish.
The AI pilot situation in 2025 looks structurally identical.
The five percent pattern is not random
McKinsey's 2025 report puts roughly two-thirds of organizations in pilot purgatory, running experiments with no production deployment. MIT research, cited in the same report, finds only five percent of AI pilot programs generate measurable profit and loss impact. Across nearly two thousand respondents, only about five and a half percent report that AI contributes more than five percent of EBIT. These numbers appear independently in separate research streams and converge on the same figure. That is not noise. It is a structural outcome produced by how pilots get selected and managed.
Melissa Perri's diagnosis is worth taking seriously here. She argues the failure is not technical. It is organizational: undefined outcomes, unclear ownership, no selection criteria, no success metrics. Teams treat AI adoption as a technology problem when it is a portfolio problem. The Fortune analysis of firms stuck in pilot purgatory adds a detail Perri's framing predicts: the pilots that accumulate without producing results are owned by enthusiasts, not enterprise leaders. No cross-functional ownership. No ninety-day prove-and-scale wave. No one with the authority or incentive to stop them.
The experimentation argument assumes the experiments are actually running
The strongest objection to narrowing your pilot portfolio is James March's. Exploration and exploitation have different learning dynamics. Exploratory initiatives are harder to measure, their value is uncertain at the time of assessment, and a value-versus-difficulty diagnostic applied before the learning happens will systematically favor use cases whose value is already legible. You end up concentrating effort on the use cases that look like winners before the experiments run, which is not the same as concentrating effort on the actual winners.
This is a real argument. It is also an argument about a different situation than the one most founders are in.
March's framework requires that the pilots generate usable signal. A pilot without defined success criteria, without cross-functional ownership, without a fixed timeline and a scale-iterate-pivot-stop decision rubric, is not generating signal. It is generating activity. The Iternal piece on pilot purgatory identifies complexity overload as a distinct failure mode: running too many pilots simultaneously prevents any single one from receiving the focused attention needed to reach a decision point. The option value March's argument is protecting does not exist in a pilot with no mechanism for producing a decision.
What a value-versus-difficulty diagnostic does and does not do
The AI whitepaper approach describes a structured process: ideation, then assessment, then prioritization, with a value-versus-ease-of-implementation matrix and clustering of use cases into a focused set of three to six per wave. This does not eliminate exploration. It eliminates the accumulation of pilots with no path to production.
The S&P 500 ambidexterity study from Helsinki and Minnesota found that balanced exploration-exploitation portfolios link to better financial performance, with the optimal allocation driven by industry R&D intensity. The firms that escaped pilot purgatory in the Fortune analysis did not stop experimenting. They scrapped most pilots, selected three to five high-impact use cases, formed cross-functional leadership squads, and ran ninety-day prove-and-scale waves. The structure preserved experimentation. It eliminated the condition where no experiment ever reaches a decision.
If you are running five or more AI pilots with no production deployments, the diagnostic question is not which pilots are most technically promising. It is which pilots have a named owner, a defined success metric, and a fixed date by which the scale-iterate-pivot-stop decision gets made. The ones without all three are not exploratory bets. They are budget consumption with an innovation label attached.
The Dutch SME thesis found no clear optimal allocation formula across its eight-year sample, which is a genuine caution against rigid rules. Three to five use cases per wave is not a formula. It is the output of applying selection criteria that most pilot portfolios currently lack entirely.

Read next

Human-Centered Transformation
AI Pilots Don't Fail at the End
Most AI pilots fail before they start. MIT, RAND, Gartner, and IDC data show why narrow scope and named ownership separate the 5% from the rest.
4 min read

Getting to ROI
Your AI Pilot Isn't Failing Because The AI Is Bad
Most AI pilots fail before the technology gets a fair test. Here's the structural fix founders miss before day one — one use case, one metric, one decision.
3 min read

The Execution Layer
Why AI Pilots Succeed But Never Reach Production
Your AI pilot worked. So why is it still a pilot? Four structural reasons enterprise AI stalls between demo and deployment, and the questions that fix it at…
3 min read