Your AI Pilot Numbers Are Probably Wrong

Brynjolfsson, Li, and Raymond tracked 5,179 customer support agents through a staggered rollout of an AI assistant and found a 13.8 percent average increase in issues resolved per hour. That number gets quoted constantly. What gets skipped is the line two paragraphs later: highly skilled workers showed minimal change, while workers with two months of tenure gained 34 percent.
If you ran a two-week AI pilot on email response time and got a clean before/after number, you probably think you measured something. You measured an average across people who respond differently to the tool by design.
The number you produce depends on who took the pilot
The Brynjolfsson study shows that new hires using AI assistance reached the output level of untreated agents with six or more months of tenure. The tool essentially transferred tacit knowledge from experienced workers into model suggestions, which is why novices captured nearly all the gain. Experienced workers already had that knowledge. The AI gave them nothing they didn't already carry.
This is not a quirk of customer support. The OECD's review of generative AI productivity research puts worker-level gains between 10 and 56 percent depending on task type and worker characteristics. That range exists because the worker matters as much as the tool. A pilot that aggregates across experience levels doesn't sit somewhere in that range. It produces a number that could mean anything within it.
What an unsegmented result actually tells you
A founder with a senior-heavy team who runs an unsegmented pilot and sees a 2 percent improvement will conclude the tool failed. The research says the tool genuinely doesn't move the needle for experienced workers, so the conclusion isn't wrong, but it's also not generalizable. Bring on two junior hires and the number changes entirely.
The inverse is the more dangerous case. A founder with a junior-heavy team sees a 30 percent improvement, attributes it to the AI subscription, and scales the tooling. The research says that gain reflects the novice-gains pattern, which may shrink as those workers accumulate experience. The pilot number looked like a product effect. It was a tenure effect.
The Research Scope section of the underlying research names this failure mode directly: reported gains might reflect novelty, selection bias, or measurement artefacts, and founders need structured diagnostics before treating AI budgets as recurring investments.
The steelman worth taking seriously
A reasonable objection holds that any structured before/after measurement beats gut feeling. A founder deciding whether to keep a $50 per month subscription doesn't need a controlled study. A rough aggregate signal is enough to kill or keep the tool.
That argument works for one decision. It breaks down the moment the number gets used for anything else: reporting to investors, justifying a team-wide rollout, or comparing tools against each other. An unsegmented number that cannot be interpreted is not more useful than no number. It produces confident conclusions that point in the wrong direction.
The fix costs almost nothing
Segmenting a pilot by experience level doesn't require 5,179 agents or a research team. It requires one column in a spreadsheet: joined in the last six months, or longer than that. Two groups. Record baseline email response time for each group separately. Run the pilot. Compare within groups, not across them.
A founder with four people on the team who records that split will know whether the 20 percent improvement came from the two junior staff or from everyone. That distinction determines whether the number will hold as the team grows, whether it will shrink as those junior staff gain experience, and whether the tool is worth expanding.
The OECD's 10 to 56 percent range, which the counterargument uses to validate unsegmented results, actually makes the case for segmentation. That range exists because worker characteristics drive the outcome. A founder who doesn't know where their team sits in that range has produced a number, not a measurement.
Record who took the pilot and how long they've been on the team. The pilot result will mean something specific instead of something plausible.

Read next

The Execution Layer
Your AI Pilot Isn't Failing Because the Tool Is Wrong
Most small business AI pilots collapse before they prove anything. Here's the structural reason why, and what a 90-day lead response pilot looks like when
5 min read

The Execution Layer
AI Pilots Don't Fail at the Demo
Most AI pilots fail before anyone writes a single line of code. Here's why the 80% failure rate traces to one missing document, and how to fix it before you
3 min read

Getting to ROI
Your AI Pilot Isn't Failing Because The AI Is Bad
Most AI pilots fail before the technology gets a fair test. Here's the structural fix founders miss before day one — one use case, one metric, one decision.
3 min read