Your Team's AI Outputs Are Bad Because of How They Practice

Your customer service rep types "write a reply to this complaint" into ChatGPT, pastes the email, and accepts whatever comes back. The output is grammatically fine and completely wrong in tone. She edits it manually for ten minutes. Nobody calls this a prompting problem.
The failure mode nobody names
Research by Zamfirescu-Pereira and colleagues at ACM CHI watched ten non-experts try to write effective prompts in real time. The participants did not fail because they lacked technical vocabulary. They failed in three specific ways: they iterated opportunistically (changing whatever felt wrong without a theory of why), they over-generalised from single successful exchanges, and they treated the AI like a conversational partner who responds to social warmth rather than structural instruction. Reading a list of prompt tips does not interrupt any of these behaviors. It gives people new vocabulary for the same old habits.
This is why I have no patience for generic prompt guides. A PDF titled "50 ChatGPT prompts for small business" is not useless because it is badly written. It is useless because the failure mode it addresses does not exist. Nobody is prompting badly because they lack a list of templates.
What task-grounded repetition does differently
Klinkenberg and colleagues ran two studies on prompt engineering as a learnable skill. Study 1 found that prompt quality explains a statistically significant portion of output quality variance across tasks. Study 2 found that brief exposure to worked examples, not abstract rules, improved participants' ability to apply targeted prompting strategies. The mechanism is example-to-example comparison: a practitioner writes a prompt, evaluates the output against a standard they already hold, and adjusts.
That last clause matters. An SMB employee who handles customer complaints daily already knows what a good reply looks like. When she writes a prompt and the output misses the tone, she does not need a rubric to tell her it missed. She knows. That prior knowledge of the expected output is what was absent in every lab study where non-expert prompting failed. The Zamfirescu-Pereira participants had no independent basis for judging whether the AI behaved correctly. Your team does.
Wang and colleagues confirmed the output side of this through regression analysis on real developer-AI interactions pulled from GitHub logs. Measurable prompt properties, including readability and structural completeness, predicted conceptual consistency, perceived usefulness, and correctness of generated outputs. The relationship between prompt quality and output quality is not theoretical.
Where this breaks
The journalist study on arXiv is the honest limit on this argument. Twenty-nine journalists received in-person, task-specific prompt training built around their actual work. Self-rated expertise improved. Accuracy gains appeared on lower-difficulty tasks. On harder tasks, the results were inconsistent, and perceived helpfulness was mixed even with a live instructor running the sessions.
A self-directed fifteen-minute daily exercise with no instructor and no external feedback loop will not outperform that. For complex analytical tasks, multi-source synthesis, or anything requiring the AI to reason across ambiguous inputs, daily practice on familiar documents is not enough. The claim here is narrower: for the tasks that fill most SMB workdays, customer email replies, routine reports, meeting summaries, the journalist study's data supports task-grounded practice rather than undermining it.
The exercise that fits inside a workday
Pick one document your team produces every day. A customer reply, a weekly summary, an intake form response. Spend fifteen minutes writing three different prompts for the same output. Compare the results against the standard you already hold. Note what the weaker prompts missed, not in abstract terms but in the specific language of that document type.
Do this with real work, not practice scenarios. The feedback loop only closes when the stakes are real enough to tell you whether the output would have actually been sent.
After two weeks, the over-generalisation failure Zamfirescu-Pereira documented starts to erode, not because the team learned new rules, but because they accumulated enough counter-examples to stop trusting any single successful prompt as proof of a working strategy.

Read next

Human-Centered Transformation
Two Staff, Not a Data Team
Most SMB founders assume AI output quality is a tool problem. The research says it's a communication problem, and it's fixable in weeks, not quarters.
3 min read

Human-Centered Transformation
Role-Specific AI Guides Work Better Than AI Training
General AI training sessions don't change how your team uses AI day-to-day. Here's why role-specific guides outperform them, and what to put in one.
3 min read

Human-Centered Transformation
Your Team's AI Training Problem Isn't YouTube
Most teams learn AI from scattered videos, leaving real gaps in prompt habits, output review, and error reporting. Here's a 2-hour session plan that fixes that.
3 min read