Why Your AI Content ROI Number Is Probably Wrong

Founders who adopt AI writing tools report feeling faster. They describe producing more posts per week, shorter drafting sessions, fewer hours on email. The feeling is real. The number they put on it is not.
The problem is not that AI fails to save time. Noy and Zhang's randomized experiment, published in Science, found that workers given ChatGPT access finished writing tasks 40% faster and received quality grades 18% higher than the control group. Dell'Acqua and colleagues ran a field experiment with 758 Boston Consulting Group consultants and found that those using GPT-4 worked 25% faster and produced outputs rated more than 40% higher in quality. These are controlled experiments, not surveys. The savings are real and measurable when a baseline exists.
The baseline is exactly what most founders skip.
The measurement problem that appears after adoption
When a founder adopts an AI writing tool under time pressure, which is when most adoptions happen, they open the tool and start using it. They do not spend the prior week timing their drafts. They do not log how long a blog post took before the tool existed. They start saving time immediately, and they start measuring immediately, and those two things are not the same as measuring the difference.
Brynjolfsson and colleagues studied 7,137 knowledge workers across 66 firms using a generative AI tool integrated into daily software. Email time savings of two hours per week only became detectable in the latter half of a six-month study, using actual software usage logs, with a pre-treatment observation period built into the design. The savings did not show up in early measurement windows. Without the pre-treatment baseline, those two hours per week would have been invisible in the data.
A founder relying on memory of how long drafts used to take is working with far less signal than 7,137 workers and software logs, and the Brynjolfsson study still required a prospective design to detect the effect.
Reconstructed baselines look credible until someone asks a follow-up question
The reasonable objection here is that founders with invoices, calendar records, or consistent task history do not need a prospective baseline. A founder who paid a freelance writer for six months before switching to AI has a paper record of pre-AI content costs. That record is a workable reconstruction.
This objection holds for outsourced content with documented costs. For founders doing their own writing with no time-tracking records, it does not hold for a specific reason the research names directly: the "workslop" problem. AI tools produce higher output volume, which makes founders feel more productive, even when per-piece quality has dropped. A founder who remembers spending four hours on a newsletter before AI and now spends ninety minutes is comparing two different tasks if the AI-assisted version is lower quality. The retrospective baseline treats those as equivalent. The inflated ROI number that results is not a lie the founder is telling deliberately. It is a structural artifact of measuring output volume without measuring output quality.
The BCG study found quality gains of more than 40% for tasks inside the AI capability frontier. It found the opposite pattern for tasks outside that frontier: consultants using GPT-4 on unsuitable tasks performed worse than the control group. The capability frontier is not labeled on the tool. Founders discover it by tracking quality, not just speed.
What a one-week baseline actually requires
The measurement is not complicated. Before using an AI tool for content, or starting from the next piece if the tool is already in use, log the time from blank page to finished draft for each piece. Log the number of pieces completed in the week. Rate the output quality on a simple scale before publishing.
Run the same log for one week with the AI tool active.
The difference between those two weeks, converted to an hourly cost, is the rough ROI figure. Noy and Zhang's experiment measured gains within single task sessions, not over months, which means a one-week window captures real signal. The caveat is that one week is also short enough to miss rework costs and prompt refinement learning curves that appear later.
The Lean Startup tradition's validated learning requirement is not about rigor for its own sake. A baseline set before the experiment is the only thing that separates a defensible ROI number from a feeling.

Read next

The Execution Layer
Time Saved Is Not Money Earned
Most founders tracking AI ROI measure the wrong thing. Here's why time logs mislead, what the research shows about task gains versus firm-level returns, and how
3 min read

The Execution Layer
Measuring AI ROI When You Can't Prove It's Working
Small business owners paid for AI tools but can't show results. Here's how to track net time saved, error rates, and task volume to calculate real ROI.
4 min read

The Execution Layer
AI Helped Is Not a Number
Most founders believe AI is working. Almost none can show a before-and-after. This is the measurement problem that compounds quietly until the budget
5 min read