Your data isn't dirty. It's silent.
7 min readEquipo de Nodu

Why you have spent two years waiting to "sort the data out" before doing anything with artificial intelligence, and why that wait probably wasn't necessary.
There is a sentence said in almost every meeting where a mid-sized company considers doing something with artificial intelligence. The finance director says it, or the IT director, or the owner, and it usually arrives around minute twenty:
"Before any of that, we need to sort the data out."
Nobody ever argues with that sentence. It is sensible, it is prudent and it sounds like professional judgement. And from what we see in the companies we talk to, it is the best-mannered excuse there is for doing nothing for two years.
The project that never starts
What happens after that sentence is always much the same. A data quality assessment is commissioned. It turns out — surprise — that there are duplicates, empty fields, customers entered three times under three different names, and a product reference somebody changed in 2019 without telling anyone. A clean-up project is proposed. The project is big, dull and produces no visible result until the very end. It gets scheduled for next quarter. And the one after that.
Meanwhile, the company carries on working exactly as before. With the same imperfect data that, curiously, does not stop it invoicing, collecting, running payroll and closing the year.
That is the clue almost nobody picks up: if your data were genuinely unusable, your company could not operate. And it operates. Every day. Which means the problem is not as big as you have been told — it is of a different nature.
Dirty is not the same as silent
Dirty data is data that is wrong: a mistyped amount, a duplicated customer, an impossible date. It exists, of course. Every company has some, and it needs correcting when it shows up.
Silent data is something else. It is perfectly correct data that cannot answer a question.
The difference shows up in any example you like. Say you want to know which customers systematically pay you late. The data exists: every invoice has its issue date, its due date and the date it was paid. None of it is dirty. All of it is exact. And yet nobody in your company can answer that question without exporting three lists into a spreadsheet and spending an afternoon cross-referencing them.
The data was fine. It just wasn't talking to itself.
Another one: you want to know the real margin on each project. The time is logged, the purchases are recorded, the invoices are issued. All correct, all in the same system. And the answer still takes a week to arrive, if it arrives.
And one more, the one we like best because it is the most common: your sales team has meeting notes, emails with customers, terms agreed over the phone and warnings along the lines of "don't call this one on a Monday". None of that is dirty. None of it is even data, in the sense a report understands. It lives in people's heads, in inboxes and in loose documents. It is silent by design.
Almost everything a small company calls a "data quality problem" is really this: correct information that nobody can interrogate.
Why we got this so wrong
The confusion has a historical cause, and it is worth understanding, because it explains why the advice keeps being repeated when it no longer applies.
For twenty years, anything you wanted to ask your data had to be programmed. Somebody wrote a query, somebody designed a report, somebody built a dashboard. And those queries are literal: they break if a field is empty, if a customer appears twice or if the format of a date changes. A query does not interpret. It executes.
In that world, "clean the data first" was excellent advice. It was, in fact, the only possible order: without homogeneous data, the report either failed or gave you a false figure, which is worse.
What has changed is that there is now something between the question and the data that does interpret. A system that can read a scanned invoice, understand that "Industrias García SL" and "IND. GARCIA, S.L." are the same company, work out from a transcript that what the customer asked for was more time to pay, and tell you that a pattern is repeating. Not because somebody programmed it for that, but because it understands what it is reading.
And that reverses the order of things. You used to have to clean up in order to ask. Now you can ask in order to find out what is worth cleaning up.
We did it in that order out of sheer impatience, and it turned out to be the right order. We put our own artificial intelligence systems to work inside our ERP without having cleaned anything up, asking things nobody had been able to ask before. The first thing that came out was not an answer: it was a map of what was broken and what did not actually matter. Of the problems on that map, half of the ones we had written down as needing fixing turned out to be irrelevant. The other half, the half that really did affect decisions, was fixed in a week — because by then we knew exactly what to fix and why.
What waiting costs
It is worth saying what the well-meant advice costs.
A company that decides to clean up before asking commits to a long, expensive project with no interim results, decided blind: since it does not know which questions it will want to ask, it cleans everything equally. It cleans fields nobody will ever look at as carefully as the ones that decide its margin. And when it finishes — if it finishes — it discovers that the questions it actually cares about needed other data, which was not in the plan.
While that goes on, the company keeps making decisions with the information at hand, which is worse than the information it could have. That is the real cost, and it appears in no budget: two years of decisions taken in half-light, waiting for a clean-up that never quite gets there.
There is an asymmetry here worth thinking about. Cleaning up without knowing what you are going to ask is expensive and slow. Asking without having cleaned up is quick and cheap — and the worst possible outcome is that you find out what is wrong. When one of the two options has a floor that low, the order decides itself.
You don't need a big project to find out
And here comes the uncomfortable part for the industry we work in: finding out whether your data is dirty or merely silent is not a project. It is an afternoon.
All it takes is asking your system three or four questions that nobody can answer today without opening a spreadsheet, and seeing what happens. If it answers them, you never had a data quality problem: you had an access problem. If it doesn't, you now know exactly where the broken part is — not in the abstract, but in the specific place that stops you knowing what you want to know.
The reassuring thing is that this test requires touching nothing. It can be done on a copy of your system, on another machine, with nothing installed on yours and no risk to what works for you today. If the answer is that you are fine as you are, you have saved yourself a two-year project. And if you are not, you have swapped a blind general clean-up for a short list of specific things.
The three questions, and what it costs to ask them
We are Nodu, an Odoo agency, and this is exactly what we do: a view of your business you can't see today, built on a clone of your ERP, at a fixed published price. We work this way because we come out of an artificial intelligence lab — AiKit Research — and we have spent two years running our own company on Odoo with our own agents working inside it. We tell that story in full in The invisible list.
Before any of that there is something that costs nothing: the X-ray of your Odoo is free. We tell you what you have, what you are not using and what can already be asked with what is there — without cleaning anything first. If it turns out you don't need anything, we will be the ones to tell you.
So don't bring us tidy data. Bring the three questions nobody in your company can answer today without a spreadsheet. Write to us and we'll try them.
That, in the end, is the only difference that matters. Not between dirty data and clean data.
Between data that stays quiet and data that answers.
