One Data Fix Beats More Data Every Time

Your CRM, your email tool, and your spreadsheet each hold a different version of the same customer name, and none of them agree. Most founders who notice this conclude they need to fix everything before AI becomes useful. That conclusion is wrong, and it costs them months.
What the research actually shows
A controlled study testing fifteen machine learning algorithms found something counterintuitive: improving one specific data quality dimension—say, removing duplicate records or standardizing how a field is represented—produced larger accuracy gains than adding more records to the training set. The study measured six dimensions independently, including uniqueness, consistent representation, and class balance. Each dimension moved performance on its own. You do not need to fix all six before the first one pays off.
A separate survey of machine learning data quality research reached the same conclusion from a different angle: improving training data quality is often more efficient than enlarging dataset size. More data fed into a messy field gives you a bigger mess. A smaller, well-cleaned field gives you a better model.
The numbers from a real cleanup
A churn prediction study using the IBM Telco Customer Churn dataset applied three preprocessing steps to customer records: imputation for missing values, one-hot encoding for categorical fields, and outlier removal. After those three targeted fixes, logistic regression accuracy reached between 78.4 and 84.6 percent. The baseline, before cleaning, was lower. The researchers did not redesign the data architecture. They fixed specific, diagnosable problems.
Those three problems—missing entries, inconsistent category labels, extreme numerical outliers—are problems your customer list almost certainly has right now. They are not exotic. They are the default state of any CRM that multiple people have touched over two years.
Why cleaning one field does not fix the process that broke it
Here is the legitimate objection to everything above. Research on AI governance argues that without consistent standards across your systems, targeted fixes do not hold. Your staff will re-enter inconsistent email formats next month. The cleaned field degrades, and the model's accuracy degrades with it. A founder who cleans one field, trains a model, and considers the job done has not addressed the process that corrupted the field in the first place.
This objection is real. The IBM Telco dataset used in the churn study is a curated benchmark. It does not re-corrupt itself between training runs. A live CRM does. Whether the accuracy gains from a benchmark cleanup transfer directly to a live small business CRM is an inference, not a settled finding. [Inference]
The governance critique fails on one specific practical point, though. It assumes a full overhaul is a realistic alternative. For a business running on limited staff, a spreadsheet, and an email marketing tool, a governance-first position is an indefinite delay. The same research that raises the scaling concern also documents the barriers: limited staff, budget constraints, and deep skepticism toward AI. A full overhaul requires resolving those conditions first. Most founders cannot do that.
The narrower path the data quality pipeline research describes is more honest: rework one collection process, such as how your team records email addresses, then feed that cleaner data into a first model. That addresses the re-corruption problem without requiring a system redesign before you have seen any results.
What to do with this
Pick one field your AI use case depends on. If you want to predict which customers are likely to churn, that field is probably customer status or last purchase date. If you want to score leads, it is probably company size or source tag. Check that field for the three problems the churn study fixed: missing values, inconsistent labels, and extreme outliers. Fix those. Train a model. Measure accuracy before and after.
You will not have a perfect CRM at the end of that process. You will have one field you trust, one model you can evaluate, and a concrete before-and-after number to show anyone who asks whether the cleanup was worth the time.

Read next

Data as a Decision Infrastructure
Clean Two Fields, Not Fifty
Most data cleanup advice assumes you have a data team. You don't. Here's what actually moves AI accuracy when you're working alone with free tools.
4 min read

Data as a Decision Infrastructure
Clean Your Customer Data Before AI Breaks It
Small businesses lose revenue to bad CRM data before they ever touch an AI tool. Here's a seven-step checklist using Google Sheets and OpenRefine.
3 min read

Data as a Decision Infrastructure
Fix Your Customer Data Before Your AI Does
Most small businesses run AI tools on customer records where fewer than half the entries are accurate. Here's what that costs and how to fix it in three days.
3 min read