Three Metrics That End Pilot Purgatory

Eighty-eight percent of AI proofs of concept stall before reaching production. The models aren't the problem. The evaluation is.
The real reason pilots drift
When a pilot has no pre-defined threshold for success, the review meeting becomes a negotiation. Someone points to the demo. Someone else says the team loved it. A third person asks for another month of data. The pilot doesn't die — it just never gets a decision. This is what IDC and consulting synthesis on AI deployment describe as the primary cause of the stall rate: not model failure, but the absence of agreed success criteria before launch.
The fix is not a twelve-metric scorecard. Wavect's twelve-metric framework and Kognitos's ninety-day evaluation model are thorough, and I think they're largely useless for the population of founders who haven't instrumented anything yet. A comprehensive scorecard that no one fills in, never gets baselined, and isn't reviewed at a fixed point produces the same drift as no scorecard at all. You need something you'll actually use before the experiment starts.
What three metrics force you to do
CloudZero's AI ROI framework maps AI value onto four buckets: revenue lift, cost takeout, velocity, and risk or quality. Time saved sits in velocity. Error rate reduction belongs to risk and quality. Net revenue or cost impact covers the financial side. Three metrics, one from each of the buckets that matter for an early-stage decision.
The point isn't the metrics themselves. The point is that picking them before the pilot forces a conversation you would otherwise skip: what does "working" mean for this specific workflow? Microsoft's six-month field experiment across sixty-six firms and more than seven thousand knowledge workers shows what rigorous pre-definition looks like in practice. Researchers instrumented Outlook telemetry, defined the unit of measurement as weekly hours spent on email, and collected a baseline before randomizing access to the AI tool. The result was a twelve to seventeen percent reduction in email time for treated workers. That number is usable precisely because the measurement method was decided before the experiment ran.
If you define "time saved" after the pilot ends, you'll find it wherever you look.
When a green dashboard still stalls
The legitimate objection to this approach is that a pilot scoring well on all three metrics still has to survive contact with a production environment. Data readiness, integration effort, auditability — none of these appear in the three-metric dashboard, and the research is explicit that pilots fail on all three. A founder who hits every threshold and then commits to scaling will discover data pipeline gaps and compliance documentation requirements only after the go decision, at higher cost.
This objection is real. It describes a different problem. The 88% stall rate is documented as an evaluation failure, not a production-readiness failure. The dashboard answers one specific question: did this workflow change in a measurable way? Data readiness gates and audit trails answer a later question, and they belong in the scaling plan, not the pilot evaluation. Conflating the two is how you end up with a governance checklist that founders ignore entirely because it asks them to solve production problems before they know if the pilot works.
What to measure and when
Set your three baselines before you touch the AI tool. For time saved, define the unit of work, measure it for two to four weeks using the same telemetry you'll use during the pilot, and write down the number. For error rate, pull the last ninety days of whatever you're trying to improve — ticket rework rates, invoice correction frequency, whatever the workflow produces when it goes wrong. For cost or revenue impact, estimate the per-unit economics of the current process, not the projected economics of the AI version.
Then set a threshold for each metric. Not a range. A number. If email processing time doesn't drop by at least X minutes per week, the pilot failed. If error rate doesn't fall by at least Y percent, the pilot failed. The threshold is what converts a measurement into a decision.
The Microsoft experiment found no significant change in meeting time or document writing time — only email. That specificity is the output of pre-committed measurement. Without it, the same experiment produces a narrative about "general productivity improvement," which is not a go/no-go signal. It's a press release.

Read next

Getting to ROI
Your AI Pilot Isn't Failing Because The AI Is Bad
Most AI pilots fail before the technology gets a fair test. Here's the structural fix founders miss before day one — one use case, one metric, one decision.
3 min read

AI as Strategy
Why Your AI Pilot Stalled (and It Is Not the Model)
67% of SME generative AI pilots never reach production. The bottleneck is not the model — it is three missing process decisions founders skip before launch.
5 min read

The Execution Layer
Why Your AI Pilot Never Ships
Most founder-led AI experiments die not from bad ideas but from missing exit rules. Here's what the research shows about why structured pilots scale.
3 min read