Data Mapping Checklist for Founders

Most founders, when asked where their customer data lives, say "the CRM" — and then pause, because they're remembering the spreadsheet their sales rep keeps, the Mailchimp list that predates the CRM, and the Stripe records that don't quite match either of them. The data isn't missing. It's everywhere.
Why full coverage fails before it starts
The academic work on SME data governance, particularly Okoro's 2021 thesis on tailored governance frameworks for small firms, documents a consistent pattern: small businesses treat data as invisible infrastructure embedded in their tools rather than as a managed asset. The failure mode isn't ignorance. It's scope. Founders who try to map everything produce a project that runs for weeks and delivers nothing. The SME governance literature is direct on this point — small firms stall on data work not because they aim too low, but because they aim too broadly.
Scoping to three categories breaks that pattern. Customer records, employee records, and financial records cover the data your business is most likely to be asked about in a compliance review, a fundraise, or an acquisition. They're also the categories where duplication causes the most visible damage.
The four-hour checklist
Start with a blank spreadsheet. Five columns: system name, data category, owner, authoritative or duplicate, and one-line quality note.
Block one hour to list every tool your business uses to store or move data. Not just databases — your CRM, your accounting software, your payroll system, your email marketing platform, your shared Google Drive folders, your inbox. Write them all down. Don't filter yet.
In the second hour, walk each system and assign it to one or more of the three categories. A tool like HubSpot touches customer records. QuickBooks or Xero touches financial records. Gusto or Deel touches employee records. Some tools touch more than one. Note that.
In the third hour, for each category, pick one system as the authoritative record. This is the system of record — the one you'd use to answer a legal question or audit request. Everything else holding the same data category is a duplicate or a feed. Mark it. You don't need to resolve the duplication today. You need to see it.
Spend the fourth hour on short conversations. Ten minutes each with the person who owns each system. Ask them one question: "If this system disagreed with another system on a customer's name or payment status, which one would you trust?" Their answer tells you whether your authoritative record is actually authoritative, or whether it's just the one you nominated.
A system-level map won't catch field-level errors
The strongest objection to this approach is a real one. Knowing that customer records live in your CRM and your marketing platform doesn't tell you whether the email field in one matches the email field in the other, whether consent flags exist in both, or which address is correct when the two systems disagree. Those are field-level questions. This checklist doesn't answer them.
Okoro's thesis is explicit that data quality management and privacy alignment require more granular work than system identification provides. The research behind this checklist agrees: deeper quality assessment takes repeated reviews beyond a first pass.
The reason to do this session anyway is that field-level mapping across all your tools, attempted on day one, is the thing that produces no output. A system-level map that gets finished and shared is more useful than a field-level map stalled in a Notion doc at week two. The first pass gives you the list of systems, the nominated authoritative records, and the visible duplication. The second pass — scheduled for next quarter — goes one level deeper into the fields that matter most.
What you do with the output
After four hours, you have a spreadsheet with every data-holding system named, every system assigned to a data category, one authoritative record per category identified, and duplicates flagged. Share it with whoever manages each system. Ask them to correct anything wrong about the ownership column.
Then schedule a thirty-minute review for ninety days out. That review is where you start asking field-level questions — specifically about the systems where your authoritative record and your duplicate disagreed during the stakeholder conversations. That's where the quality work begins.
The map you built today is not a finished governance program. Okoro's framework treats governance as evolutionary, not static. What you built is the precondition — the thing that has to exist before any of the structured work can start.

Read next

Data as a Decision Infrastructure
Fix Your Customer Data Before Your AI Does
Most small businesses run AI tools on customer records where fewer than half the entries are accurate. Here's what that costs and how to fix it in three days.
3 min read

Data as a Decision Infrastructure
AI-Ready Customer Data Checklist for Small Businesses
Small businesses connecting AI to unaudited CRM, email, and invoicing data risk confident wrong answers. Here's a 30-day checklist to fix that first.
3 min read

Data as a Decision Infrastructure
Clean Your Customer Data Before AI Breaks It
Small businesses lose revenue to bad CRM data before they ever touch an AI tool. Here's a seven-step checklist using Google Sheets and OpenRefine.
3 min read