Run the Test Before You Build the Case

Most founders trying to justify AI investment do the same thing: they find a Forrester report, locate the productivity lift percentage, and build a slide around it. The number looks credible. The board asks one question — "does that apply to our work?" — and the slide dies.
The problem is not the Forrester number. Forrester's ROI framing is real, and the controlled studies behind it are serious. Microsoft, Harvard, MIT, and GitHub have each run structured comparisons showing productivity lifts in the range of 10 to 50 percent for well-chosen tasks. The problem is "well-chosen." That qualifier does more work than the percentage. AI helps some tasks and harms others in ways that are not predictable from vendor claims or industry benchmarks. Researchers call this the jagged frontier. A founder who borrows a benchmark without testing their own tasks has not found evidence. They have found a number that sounds like evidence.
What a business case built before deployment actually is
Before deployment, a founder building an AI business case is projecting gains from process redesign that has not happened yet, on tasks whose AI performance has not been tested, in an organization whose adaptation curve is unknown. The projection might be correct. It is not defensible, because it cannot be falsified by anything that has already occurred. A skeptical board member does not need to disprove it. They only need to ask what it is based on.
The parallel test answers that question with data from your own work.
How the five-day test works
The structure Skycrumbs recommends is direct. Take one real task your team runs repeatedly — customer support triage, contract summarization, first-draft proposals, whatever generates volume. Split it. Your human team completes version A using their normal process. The AI completes version B using the same inputs. You measure three things: time to complete, correctness checked against a known standard, and output quality rated by someone who knows the work.
Five days is enough to generate a pattern on a high-volume task. It is not enough to capture what happens after your team has adapted workflows around the AI. That ceiling is real, and it matters. A five-day test measures the tool before the organization has changed to accommodate it, which means the productivity number you find is probably lower than the number you would find after six months of genuine integration. The compounding gains from workflow redesign — the ones that drive the upper end of the 10 to 50 percent range — do not show up in five days.
This is the strongest objection to the method, and it deserves to be left standing rather than explained away. A five-day test produces a conservative estimate. It does not produce a complete picture of long-term AI value.
Five days measures the tool, not the transformation
The objection is methodologically sound and practically irrelevant for a founder trying to justify an investment decision before deployment. The compounding-gains argument says the parallel test understates eventual ROI. It does not say a pre-deployment business case overstates it less. A long-horizon projection built before deployment cannot account for the jagged frontier problem, because the founder does not yet know which tasks fall on which side. The five-day test at least identifies which task types show positive results before the commitment is made. The pre-deployment business case identifies none of them.
A conservative estimate that survives board scrutiny is more useful than a speculative projection that does not. The parallel test produces the former.
What to measure and how to score it
Time is the easiest dimension. Log start and end for each task, both human and AI versions. Keep the inputs identical.
Correctness requires a standard. Before the test starts, define what a correct output looks like for your task type. For contract summarization, correctness means all material obligations are present and none are fabricated. For customer support triage, correctness means the ticket lands in the right category. Without a pre-defined standard, correctness scoring becomes subjective after the fact, which is exactly the kind of measurement trap the research warns against.
Output quality is the hardest dimension and the most important one for founders whose work product goes to clients or partners. Rate quality blind — the evaluator should not know whether a given output came from a human or the AI. Blind rating removes the tendency to forgive AI errors that would not be forgiven in human work, and it removes the reverse tendency to grade AI outputs harshly because they feel unfamiliar.
What the results tell you
A fifteen percent time reduction on a task type is not evidence of a fifteen percent long-term productivity lift. It is evidence that on that task, in that week, before process redesign, the AI completed the work faster. That is still useful. It tells you where to invest further testing, which workflows to redesign first, and which tasks to leave alone because the AI underperformed or introduced errors the human version did not.
The jagged frontier means some of your results will be negative. AI will be slower, or less accurate, or produce output your evaluators rate below the human version. Those results are not failures of the test. They are the test working. An organization-level business case built before deployment cannot surface those failures because it has no mechanism to find them. The parallel test finds them in five days, before they become expensive.
The number you bring to your board after running this test is not borrowed from a Microsoft study. It is from your work, your data, your evaluators. That is the difference between a slide that survives one question and evidence that holds up to ten.

Read next

AI as Strategy
Why Your AI Business Case Fails Finance Review
Most AI business cases are rejected because they omit change costs and name no owner. Here is the finance-literate structure CFOs need to approve AI investment.
3 min read

Human-Centered Transformation
Start with One Task, Not a System
Most founders stall on AI before they begin. Here's how to run a real experiment this week using email, a spreadsheet, or your CRM — no tech team needed.
3 min read

AI as Strategy
AI Proof of Concept = Proof of Confusion
Most AI pilots prove the technology works — not that your organisation needs it. Here's how to run strategic tests that answer the right question before you…
4 min read