When Your Reports Disagree, the Problem Starts at the Door

Your sales dashboard shows one number. Your finance team shows another. Both pulled from systems you pay for, both updated this week. The instinct is to blame whoever built the reports. The actual problem is older and sits further upstream.
The error was already there when the data arrived
SAS survey data shows 44.7% of SMBs operate with fragmented data, and the research traces that fragmentation directly to undetected errors accumulating at entry points and early pipeline stages. IBM research connects scattered, error-prone data to wrong decisions and revenue loss — not as a theoretical risk, but as a documented outcome for firms that never placed a check at the moment data entered their systems.
The errors are not exotic. A customer ID entered twice with a slightly different format. A required revenue field left blank because the form allowed it. A date recorded as text instead of a date value, which breaks every downstream calculation silently. These are not the problems that require a data scientist to find. They are the problems a short SQL rule or a no-code validation check catches before the record moves anywhere.
What the platform already does, and where it stops
Here is the objection worth taking seriously: modern SaaS tools already enforce field-level constraints. Your CRM rejects an email without an @ symbol. Your billing platform flags duplicate invoice numbers. If the platforms handle this, why write rules at all?
The answer is geography. Platform-native checks protect data inside a single tool. They do nothing for data moving between your CRM, your accounting software, and the spreadsheet your operations team maintains. A customer ID that exists in your CRM but was entered differently in your billing system passes every native check in both tools and then breaks every join you run between them. The governance research is clear on this: cross-tool fragmentation is where undetected errors live, and no individual platform's validation layer covers that boundary.
The counterargument does not collapse entirely. If your entire operation runs inside a single integrated platform, the case for manual rule-building weakens. But 44.7% of SMBs with fragmented data suggests most founders are not in that situation.
The check that actually runs
Non-specialist teams write these rules. The applied data engineering literature describes short SQL queries that flag missing required fields, duplicate records on a unique identifier, values outside an expected numeric range, and format mismatches on fields like dates or phone numbers. No-code tools like Talend and Great Expectations offer the same checks without writing code at all. The intervention point the research identifies as highest-leverage is the earliest one: data entry and the first stage of any internal pipeline, before a bad record propagates into every report that touches it.
A format check on a revenue field does not resolve a disagreement between your sales team and your finance team about when revenue counts. That disagreement is definitional, not technical, and no validation rule fixes it. This is a real limit. The thesis here is narrower: the technical errors — missing fields, duplicates, broken identifiers — that accumulate because no one placed a check at entry are the documented primary driver of the fragmentation SAS and IBM describe. Fixing the definitional problem matters too, but it requires a conversation, not a rule. The technical errors require a rule, and you do not need a data scientist to write one.
Where to place the first check
Start at the point where external data enters your systems: a form submission, a CSV import, an API feed from another tool. Add a check confirming every required field is populated. Add a check confirming no record shares a unique identifier with an existing one. Add a range check on any numeric field where an outlier would change a decision.
Run the check before the record moves downstream. Log every failure. Review the log weekly.
Your CRM, your accounting tool, and your inbox each hold a different count of the same customer right now, and none of them agree. That is not a data science problem. It is an entry-point problem with a documented, low-cost fix.

Read next

Data Foundations
Three Revenue Numbers And One Root Cause
Marketing, finance, and ops report different revenue from the same business. This is a data definition problem — and it will break any AI model you build on…
3 min read

Data as a Decision Infrastructure
Three Questions Your Data Must Answer Before AI Helps You
A practical data health check for SMB founders — no tools required. Find out if fragmented data is silently breaking your AI and analytics outputs.
3 min read

Data as a Decision Infrastructure
Your AI Won't Fix Data You Haven't Fixed First
Scattered spreadsheets and email threads don't just slow down AI projects — they cause them to fail. Here's how to migrate your most critical records into one
3 min read