How to Know if Your AI Tool Is Actually Working

McKinsey's global AI survey has tracked enterprise adoption for years and the number that keeps appearing is this: roughly 37% of organizations report some positive EBIT impact from AI, and only 6% qualify as high performers who derive meaningful financial returns. Adoption keeps growing. The 6% figure stays flat. Something is wrong with how organizations run these experiments, not with the tools themselves.
MIT's Project NANDA looked directly at enterprise AI pilots and found that 95% of them return zero measurable value. The report names two causes: weak integration and missing business context. The second cause is the one founders control from day one.
The failure happens before you open the tool
Hill-Wilson's analysis of enterprise pilot design, drawing on Google's internal KPI guidance, identifies the structural problem precisely. Most pilots fail because the organization never agrees on pilot type, primary objective, success definition, or how learning gets captured. Without those four things, the pilot produces impressions, not evidence. Workers feel faster. Managers sense improvement. No one can say by how much, compared to what, or whether it held across the full period.
This is not a technology problem. It is a design problem. And it is fixable before you run a single prompt.
Pick the right task first
Dell'Acqua and colleagues ran a field experiment with consultants at Boston Consulting Group and found something founders need to sit with. For creative product innovation tasks, consultants with GPT-4 access completed more tasks, worked roughly 25% faster, and delivered output rated at least 30% higher in quality by external graders. For a complex brand strategy case requiring multi-source integration, AI users produced more polished recommendations but reached correct solutions less often than consultants working without AI. The tool made them sound better while making them less accurate.
The implication is not that AI is unreliable. The implication is that task selection is the decision. Pick a task where AI's output is directly assessable and the quality criteria are unambiguous. Proposal drafting works because a proposal either advances to the next stage or it does not. Customer support ticket resolution works because resolution rate and handling time are already tracked. Tasks that require subtle judgment calls across multiple sources, the kind Dell'Acqua's brand strategy case represents, are the wrong starting point for a pilot.
Noy and Zhang's preregistered experiment on occupation-specific writing tasks gives the clearest benchmark for what good task selection produces. When college-educated professionals used ChatGPT on bounded writing tasks with clear quality criteria, time per task dropped roughly 40% and output quality rose roughly 18%, with lower-ability workers gaining more than experienced ones. The gains were measurable because the task was measurable.
Set the baseline before you touch the tool
The baseline is the only thing that makes day 90 mean anything. Before the pilot starts, measure the current state of the task you selected. For proposal drafting, that means time per draft, revision cycles before approval, and win rate on proposals sent. For support tickets, that means issues resolved per hour and customer sentiment scores. These numbers need to come from the four to six weeks before the pilot begins, not from memory.
Brynjolfsson, Li, and Raymond's field experiment on a generative AI assistant for more than 5,000 customer support agents shows what a clean baseline produces. The study measured issues resolved per hour before and after AI access. The result was a 14-15% average productivity improvement, with roughly one-third improvement for novice and low-skilled workers. That finding is credible because the researchers had a pre-existing baseline. Without it, a 14% improvement is invisible inside normal week-to-week variation.
Your baseline does not need to be perfect. It needs to be consistent. The same metric, measured the same way, for the same task, across the same time window.
The 90-day window is not arbitrary
The steelman against a fixed 90-day window is worth taking seriously. MIT Project NANDA explicitly names weak integration as a cause of pilot failure, and integration takes time. Workers are still building habits at week four. Prompting patterns are still evolving at week eight. A go/no-go call made on incomplete behavioral data is not rigorous — it is just early.
The counterargument holds for enterprise-wide, cross-functional deployments that require workflow redesign across multiple teams. It does not hold for a bounded, task-level pilot on a single repeatable activity.
Brynjolfsson, Li, and Raymond's study is a real enterprise deployment across more than 5,000 agents, not a lab setting. The productivity gains appeared during the staggered rollout period itself, not after a multi-month ramp. On high-volume, repeatable tasks, the signal surfaces faster than critics of short windows assume. A support queue generating 200 or more interactions per week produces enough data within 90 days to separate signal from noise. Proposal drafting at a founder-stage company may generate 20-50 proposals over the same period. That volume is sufficient for a directional finding, though not a statistically precise one.
The 90-day window also forces a discipline that open-ended pilots never develop. When the end date is fixed, you define success before you start. You do not move the goalposts when the numbers come in flat.
What the comparison actually tells you
At day 90, you have two sets of numbers: baseline and pilot period. The comparison produces one of three findings.
The metrics improved by a margin that exceeds normal variation. The tool is earning its cost. Scale it or extend the pilot to confirm.
The metrics are flat or marginally better. The tool is not producing a detectable effect on this task with this team in this window. That is a finding. It does not mean the tool is useless everywhere. It means this specific deployment did not work, and you now have evidence to say so.
The metrics got worse. This happens. Dell'Acqua's brand strategy case is the clearest example from the research. A polished output that reaches the wrong answer is worse than an unpolished one that is correct. If your proposals are getting longer and losing more deals, that is the pilot telling you something.
None of these outcomes is a failure of the experiment. All three are the experiment working. The 95% of pilots that return zero measurable value never reach this point because they never defined what measurable value would look like. You will have defined it on day one.

Read next

The Execution Layer
Your AI Pilot Isn't Failing Because the Tool Is Wrong
Most small business AI pilots collapse before they prove anything. Here's the structural reason why, and what a 90-day lead response pilot looks like when
5 min read

The Execution Layer
AI Pilots Don't Fail at the Demo
Most AI pilots fail before anyone writes a single line of code. Here's why the 80% failure rate traces to one missing document, and how to fix it before you
3 min read

The Execution Layer
AI Pilots Don't Fail Because the Model Is Wrong
Most AI pilots produce no measurable business result — not because the AI underperforms, but because founders never defined what "performing" meant.
3 min read