Before Your AI Pilot, Know What Feeds It

Only 34% of organisations hold complete knowledge of where their data is stored. If you are a founder about to run an AI pilot, that number describes your starting position more accurately than you want to admit.
The classification problem no one solves first
Here is the argument you will hear from people who think pre-pilot governance is wasted time: most founder-led pilots target narrow internal use cases — a document classifier, a churn model, a support router — and those systems sit below the EU AI Act's high-risk threshold. If the Act's strict data governance obligations only attach to high-risk systems, why run a full inventory before you know whether the rules apply?
The argument is reasonable. It also depends on a condition you cannot satisfy without doing the work it claims to skip.
To classify your pilot correctly, you need to know whether the training data contains special-category personal data, biometric proxies, or employment-linked records. Any of those push a system into high-risk territory under the Act. Small businesses average 242 SaaS tools in their portfolio, with a growing share not managed by any central function. Your CRM, your support inbox, your HR platform, and your product analytics tool each hold different slices of the same customer or employee record, and none of them agree on what they contain. You cannot answer the classification question without first knowing which systems feed the pilot. The inventory is not optional overhead you do after confirming the rules apply. It is the only way to confirm whether they apply.
What the EU AI Act actually requires before training begins
The Act's data governance obligations for high-risk systems are specific. Providers must document data origin, data preparation operations including annotation and cleaning, assumptions about what the data measures, and whether datasets reflect the geographical, contextual, and behavioural setting in which the system will operate. Teams must examine datasets for biases affecting health, safety, or fundamental rights, and document measures taken to detect and correct those biases. Automatic logs must record input data, reference databases, use periods, and the identities of people who verified results. None of this is retroactively reconstructable from a model that has already been trained.
Post-hoc audits do not fix this. A cloud provider's security certifications do not cover it. AWS and Google Cloud handle infrastructure compliance. They do not document where your training data originated, who governed it, or whether it was representative of the population your model will affect. That work belongs to you.
The checklist that takes less than a day
Start with system mapping. List every tool that stores data your pilot will touch: production databases, CRM records, support logs, spreadsheets, third-party APIs. For each one, write down what data type it holds, whether it contains personal data, and the last date someone verified the record was accurate.
Assign one named owner per system. Not a team. A person. The EU AI Act's documentation obligations require someone who verified results and governed the dataset. "The engineering team" is not an answer that survives a regulatory inquiry. Academic research on SMEs consistently shows that leaders rarely assign explicit data ownership roles, and staff focus on keeping systems running rather than defining access and retention rules. That pattern is the single fastest way to fail a post-incident review.
Flag sensitive fields before the pilot touches them. Special-category data under GDPR — health, ethnicity, religion, biometric data, trade union membership — carries stricter processing conditions. Fields that function as proxies for these categories (postcode combined with demographic data, for instance) warrant the same treatment. Tag them in your inventory. Note whether the pilot needs them at all. If it does not, exclude them from the training dataset before ingestion, not after.
When low-risk classification requires the work you skipped to prove it
The counterargument's legitimate point — that enterprise-grade governance frameworks are too heavy for SMEs — is accurate. Most published frameworks were built for organisations with dedicated data teams and do not adapt to the constraints of a five-person startup. The checklist above is not an enterprise framework. It is the minimum work required to answer the question the counterargument depends on.
2,401 confirmed data breaches in 2024 affected roughly 819 million individuals, at an average cost of $4.45 million per breach. SMEs face proportionally greater compliance pressure relative to their resources, because the regulatory obligations that apply to large firms apply equally to smaller ones. A pilot that begins as a document classifier and later connects to HR records crosses into high-risk territory. Without a prior inventory, there is no documented baseline showing what the model was trained on before the scope changed.
Run the inventory. Assign the owners. Tag the fields. Then classify the pilot.

Read next

Data as a Decision Infrastructure
AI Governance Checklist for Founders
Four steps to document your AI tools, assign data owners, and update privacy notices before a regulator does it for you. Compliant in under four hours.
3 min read

Data as a Decision Infrastructure
Ten Data Rules Before Your First AI Model Ships
59% of early AI adopters struggle to enforce data governance. Here's a ten-point policy built for founders without compliance staff.
5 min read

Data as a Decision Infrastructure
Your AI Tool Is Only as Good as Your Worst Spreadsheet
A 90-minute audit tells you more about whether AI will work on your business data than any tool comparison ever will. Here's the map.
3 min read