Why Your AI Tool Went Quiet After Week Six

A company buys an AI tool, runs a pilot, gets excited about the demo, and then three months later the tab is still open but nobody's clicking it. No one canceled the subscription. No one wrote a post-mortem. The tool just drifted to the edges of the workday and stayed there.
This is not a technology problem.
The number that should end the debate about model quality
MIT Project NANDA tracked generative AI pilots across organizations and found that 95% produced no measurable impact on the income statement. Not "modest impact." Not "hard to attribute." Zero movement on the P&L, despite real investment and real deployment. McKinsey's surveys find that only 39% of firms using AI report any earnings impact at all, and the group seeing meaningful financial returns sits at roughly 6% of all AI-using organizations.
The models these companies used were not broken. GPT-4, Claude, Gemini — these are the same tools the 6% used. The difference was not the software.
RAND Corporation's cross-project analysis offers one explanation: over 80% of AI project failures trace back to leadership and process errors rather than technical faults. A reasonable reading of that finding is that the problem sits above the workflow level — executives set wrong goals, data pipelines are a mess, nobody owns the outcome. Fix those things and the deployment works.
That argument is correct about pre-deployment failures. It does not explain what MIT documented.
When fixing leadership still leaves 95% of pilots off the income statement
The MIT figure covers pilots that were built, deployed, and running. These projects cleared the leadership and data thresholds well enough to reach production. Someone owned them. Someone signed off on the data. And they still produced nothing the finance team could find.
The RAND explanation and the MIT finding are not competing accounts of the same failure. They describe two different failure modes. RAND describes why projects die before producing anything. MIT describes why projects that produce something never show up in earnings. Both are real. The thesis that bad leadership drives all AI failure stops working the moment you look at the deployed population.
What MIT's number points to is structural. When you drop an AI tool into an existing workflow without changing how decisions get made, who owns which step, or what you measure, the tool becomes optional. People use it when it's convenient and skip it when it's not. No one notices either way because no metric captures the difference. The tool fades. The subscription renews. The failure is invisible until someone asks why the ROI slide is still blank.
What a redesigned workflow actually looks like
Invoice processing is a useful case because it's specific enough to follow and common enough to generalize from. The manual version of this process typically runs through three or four people: someone receives the invoice, someone checks it against a purchase order, someone codes it to the right account, someone approves it. Each handoff is a delay. Each person applies slightly different judgment to edge cases. The error rate compounds across steps.
The bolt-on AI version of this process adds a tool to one step, usually the data extraction step, and leaves everything else alone. The tool reads the invoice and pulls out the vendor name, amount, and date. A human still checks it. A human still codes it. A human still approves it. The process takes slightly less time on one step and the same amount of time on every other step. The savings are real but small, and because no one changed the success metrics, no one measures whether the savings are real at all.
The redesigned version starts from a different question: which decisions in this process require human judgment, and which ones are just humans doing pattern-matching that a model does faster? The answer, for a standard three-way match between invoice, purchase order, and receipt, is that most of it is pattern-matching. The redesign routes matched invoices directly to payment queue without human review. A human sees only the exceptions — mismatches, missing POs, amounts outside tolerance. The AP clerk's job shifts from processing every invoice to owning the exception logic: writing the rules, reviewing the flagged cases, and deciding when a pattern of exceptions signals a vendor problem worth escalating.
That shift requires two things most bolt-on deployments skip. First, you need a metric that tracks exception rate over time, not just processing speed. If the model flags 40% of invoices as exceptions in month one and 12% in month six, that's evidence the exception rules are improving. If it's still 38% in month six, something is wrong with the model's training data or the rules themselves. Without that metric, you never know which one it is. Second, you need someone whose job description includes owning that number. Not "uses the AI tool" — owns the exception rate. That person has authority to change the rules, escalate vendor issues, and request model retraining when accuracy degrades.
What ownership actually means in practice
Most AI deployments assign a tool owner, not a workflow owner. The tool owner makes sure the software is running and users have logins. The workflow owner is responsible for whether the process produces the right outputs at the right cost. These are different jobs. In a redesigned workflow, the workflow owner holds the success metric — in the invoice case, something like cost per invoice processed, exception rate, and days payable outstanding — and has the authority to change how the AI is used when those numbers move in the wrong direction.
S&P Global and Gartner data shows fewer than half of AI projects that enter development reach production deployment. The ones that do reach production and still fail tend to share one characteristic: the metric that would reveal the failure either doesn't exist or belongs to no one. The invoice automation deployments that work are not technically superior to the ones that don't. They have a number on a dashboard that someone checks every week and a person whose performance review includes that number.
That's not a technology insight. It's closer to an accounting insight. You cannot manage what you do not measure, and you cannot measure what no one owns. The 95% of generative AI pilots that left no trace on the income statement were not invisible because the models failed. They were invisible because no one built the instrument to see them.

Read next

The Execution Layer
AI Workflows Beat AI Tools Every Time
Most AI pilots die between demo and deployment. The reason isn't the model—it's that no one wrote down who owns what, when it runs, or what it's supposed to
5 min read

Human-Centered Transformation
AI Change Management: Why Your Rollout Stalled
Your AI pilot succeeded. Your rollout didn't. The gap isn't the model — it's the absence of role-specific work redesign, manager engagement, and structured…
4 min read

The Execution Layer
Why AI Pilots Succeed But Never Reach Production
Your AI pilot worked. So why is it still a pilot? Four structural reasons enterprise AI stalls between demo and deployment, and the questions that fix it at…
3 min read