"Our data isn't perfect. Can we still start with AI?" That is the most common question we get in a first meeting. The reassuring answer: your data doesn't have to be perfect. The less simple answer: it does have to be good enough for the specific application. The quality requirements for an internal search assistant differ fundamentally from those for assessing financial transactions. The right question isn't whether your data is perfect, but whether the relevant data is reliable enough for this task, at an acceptable risk.
Start with a concrete application
A general assessment of your data quality runs on forever and rarely produces a usable conclusion. You are better off starting with one narrowly defined use case: searching technical documentation, extracting purchase invoices automatically, building sales forecasts, classifying complaints, drafting quotes or detecting production anomalies. Only once that application is sharply defined can you determine which data you need and which quality requirements apply. That way you tie data questions to a concrete business outcome instead of an abstract ambition.
Seven questions to assess your data
Once you have chosen an application, the seven questions below help you quickly gauge whether your data can support it. We ask the same questions in every first analysis meeting with a client.
Is the necessary data available?
In many companies, the relevant information is scattered across mailboxes, shared folders, Excel files and legacy software. Only when that information is accessible in a controlled way can an AI application actually work with it. Availability is not just about existence, it is also about reachability.
Is the information complete enough?
Missing fields aren't always a problem. They become a problem only when those specific fields are needed for the application. A payment forecast with missing due dates has a big impact on the result. A summary of internal procedures survives a few small gaps without much damage.
Is the data correct?
Wrong product codes, duplicate customer records and outdated prices inevitably lead to wrong results. AI doesn't magically make bad data reliable. What AI does do: process errors faster and at greater scale. Which is exactly why you look at the source data first.
Does everyone use the same definitions?
What exactly is an "active customer"? When do you call an order "late"? Which costs belong in which category? As soon as departments use different definitions, an AI system produces contradictory conclusions from the same underlying data. Clarifying definitions is usually cheaper than straightening out the consequences later.
Is the data current enough?
An AI assistant that serves you manuals from three years ago gives technically correct but operationally outdated instructions. So define how new, changed and withdrawn information gets into the system, and how quickly.
Is the origin traceable?
For important decisions it must be clear where the information came from. Can the user click back to the source in one step, or does the output remain a black box you have to take on faith? Traceability isn't a luxury, it determines whether people dare to use the system.
Are you allowed to use that data for this purpose?
Not all information can just go to any AI platform. Check confidentiality, personal data, contractual agreements, intellectual property, access rights, retention periods and security requirements before you upload a dataset to an external service. This question belongs in the first selection meetings, not in the last week before go-live.
Not every application demands the same data quality
The quality bar depends on what the application does with the output. Broadly, AI applications fall into three categories, each with its own standard.
Supporting applications
Think searching, summarising or drafting a first version of a document. An employee reads the result and corrects where needed. Some incompleteness is acceptable as long as the sources remain visible and the user can verify quickly.
Operational applications
Here we run processes with them: classifying tickets, planning, bookings. The data has to be more consistent and the system must handle exceptions clearly, because there is less human control between input and action.
Decisive applications
Financial, commercial, HR or safety decisions based on AI output demand high standards: correctness, explainability, control, security and a solid audit trail. You don't start here without a thorough risk analysis and a clear plan for human intervention.
How do you recognise insufficient data quality?
There are clear signals that your data setup isn't ready for certain applications yet. Think shadow lists in Excel next to the official systems, reports that give different KPI numbers for the same question, manual corrections on product or customer data, no clarity on which version is current, the same information in multiple systems in parallel, important fields that staff fill in freely, exceptions that live only in email threads and history you can barely find. These signals don't mean AI is impossible. They do mean the risk goes up and you need to plan extra work before you start a use case.
You don't have to clean up all data first
The classic pitfall is waiting until everything is perfect. That moment rarely comes, and in the meantime you fall behind companies that do move. So choose a focused approach: determine which data the use case really needs, measure the current quality of exactly that data, identify the errors with the biggest impact, only fix what is needed for this application, test with realistic examples and monitor quality during use. That way you tie clean-up to a concrete business outcome instead of an endless housekeeping exercise no one ever finishes.
Test with hard examples
You must not test an AI application only with tidy, complete and carefully chosen data. Also use incomplete documents, unusual formats, duplicates, old and new versions side by side, rare exceptions, faulty input and contradictory information from different sources. Only then do you see how the application holds up in the daily reality of your company, where nothing is as clean as in a demo.
Measure data quality in business impact
Technical percentages on completeness and correctness are useful, but the impact on your business counts more. An error in an internal search result is a different weight class from an error in a payment proposal or a production instruction. So assess how often errors occur, how quickly you spot them, how much correction work they cause, what financial or operational damage they do and whether an employee can intervene in time before the error works its way through the process.
Conclusion
Your data doesn't have to be perfect, but it does have to be suitable for the application and the associated risk. Choose one concrete use case, determine which data is critical for the result, and test whether that data is available, correct, current and traceable. Good data isn't a goal in itself. It is the condition for a reliable business outcome with AI.
Want to know if your data is ready for the first AI application?
The SEMANU Analysis tests one concrete AI use case against your actual data: availability, quality, definitions and origin. No general score, but a go/no-go per application. One-off investment from 4,400 euro.
Frequently asked questions
Does my business data have to be perfect before I can start with AI?
No, but it has to be good enough for the specific application. A search assistant tolerates incomplete information as long as the source stays visible. An AI that assesses financial transactions demands correctness, explainability and an audit trail.
How do you test whether AI works with bad data?
Don't only give the system tidy examples. Test with incomplete documents, unusual formats, duplicates, old versions, rare exceptions and contradictory information. Only then do you see whether the application holds up in daily practice.
Which data questions should you ask for an AI use case?
Is the data available and accessible, is it complete enough for this purpose, is it correct, do departments use consistent definitions, is it current, can you verify the origin, and are you contractually and legally allowed to use it for this purpose.
Do you have to clean up all business data first for an AI project?
That is the classic pitfall, because that moment rarely comes. Pick one concrete use case, measure the quality of exactly that data, only improve what has the most impact and test during use. That way you tie clean-up to a concrete business outcome.