Writing
Master Data Is the Real AI Readiness Test
You do not have a model problem. You have four customer records for one customer.
Image One customer, four records
painterly editorial illustration, four nearly identical index cards pinned in a row on a dark board, each slightly different in wear and hand, deep navy surfaces, warm amber lamplight, soft teal shadow beneath the cards, tactile paper and brass pins, quiet unease, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw Most AI readiness assessments measure the wrong things. They look at cloud posture, data platform maturity, available skills, and executive sponsorship. All of that matters. None of it is usually what stops the program.
The thing that stops the program is that the organization does not agree with itself about what a customer is.
Master data is the binding constraint on enterprise AI, and unlike model capability it is entirely within your control, measurable today, and boring enough that nobody has volunteered to own it.
Why this bites agents harder than it bit reports
Enterprises have lived with imperfect master data for decades. It is survivable because humans absorb it silently.
A person pulling a report sees three rows for what is obviously one supplier and mentally merges them. A person entering an order recognizes that the customer with the trailing period in the name is the same account as the one without. A planner knows that the item flagged as active has not been purchased in four years and skips it. None of that judgment is written down anywhere. It lives in people, it is applied thousands of times a day without comment, and it is invisible until it is absent.
An agent does not have it. An agent reads what is there, at face value, at speed, and acts. The tolerance the organization thought it had was never tolerance. It was unpaid human correction, and automation is the moment the invoice for it arrives.
This is why programs get blindsided. The data was "good enough" for twenty years of reporting, so nobody classed it as a risk. It was good enough because the people were carrying it.
The five question test
You can assess this in an afternoon, without a platform, without a vendor, and long before you have chosen a model.
Can you name the system of record for each core entity? Customer, supplier, item, employee, chart of accounts. One name each, said quickly, without a debate breaking out. If three systems create customers and reconciliation runs nightly between them, you have three opinions and a schedule, not a system of record.
Figure The readiness test
Draw a clean editorial checklist diagram titled The five question master data readiness test. Five numbered rows, each with a question on the left and a pass condition on the right. Row one, Can you name the system of record for each core entity, pass condition One name, no debate. Row two, Do duplicates have an owner, pass condition A named person, not a committee. Row three, Is there a definition of the entity in writing, pass condition Written, current, and used in disputes. Row four, Can a wrong value be traced to who set it, pass condition Yes, within minutes. Row five, Does anyone measure quality on a schedule, pass condition A recurring number somebody reads. At the bottom a single band reading, Three or more failures means the model is not your constraint. Style, restrained editorial infographic, deep navy and slate on a warm off white ground, one amber accent on the bottom band, thin rules, generous whitespace, sans serif labels, no icons, no gradients, no clutter. Do duplicates have an owner? Not a committee, not a backlog, a person whose job includes it. Every large enterprise has duplicate master records. The healthy ones have somebody whose week includes merging them, with the authority to decide which survives.
Is the entity defined in writing? What makes a customer one customer. A legal entity, a ship-to, a billing relationship, a parent group. Most organizations have never written this down, which is precisely why they have four records: four teams each answered it differently and each was locally correct.
Can a wrong value be traced to who set it? Given a bad credit limit or a wrong tax group, can somebody find out who changed it, when, and from what, within minutes. If the answer is a database restore, an agent that touches this data cannot be governed, because you will be unable to distinguish its mistakes from anyone else's.
Does anybody measure quality on a schedule? A recurring number that a named person actually reads. Duplicate rate, completeness on required fields, records untouched for a suspicious length of time. Unmeasured quality does not stay level, it drifts downward, because every process has an incentive to create rather than to reconcile.
Three or more failures and the model is not your constraint. Fix the master data first, and the honest version of that sentence is that fixing it is a multi-quarter program with an owner and a budget, not a workstream inside an AI initiative.
Why the cleanup keeps getting deferred
It is not that nobody knows. Everybody knows. It gets deferred for structural reasons worth naming, because naming them is how a director gets the work funded.
The cost is borne centrally and the benefit lands everywhere else. Nobody's quarterly objective improves visibly because duplicates fell. That makes it perpetually the second priority of many people and the first priority of none.
It has no end state, which makes it hostile to project funding. Master data quality is a maintained property, like cleanliness. Programs that treat it as a one-time cleanup deliver a genuine improvement and then watch it decay, which teaches the organization that the work does not stick, which makes the next cleanup harder to fund.
And it is politically expensive in a way technical work usually is not. Deciding what a customer is means telling several teams that their local definition, which works fine for them, is no longer the one that counts. That is an organizational decision wearing a data costume, and it needs somebody with the standing to make it.
What to do instead of a boil-the-ocean program
Scope the cleanup to the agent's actual reach. You do not need enterprise-wide perfection. You need the entities that this use case reads and writes to be trustworthy, which is a far smaller and far more fundable problem. Let the first project pay for the first slice.
Fix creation before you fix history. Cleaning duplicates while the process that generates them runs untouched is a treadmill. Validation at entry, a required search-before-create, and a single owning path are less satisfying than a big remediation number and worth considerably more.
Instrument before remediating. Publish the duplicate rate and the completeness numbers before you improve anything. It gives you a baseline, it makes the improvement visible to people who fund things, and it is the only way to notice the decay later.
Let the agent find the mess. This is the genuinely good news. Reconciliation, duplicate detection, and completeness checking are Tier 1 work: nothing is committed, a wrong answer costs somebody's attention, and the tolerance for error is high because the alternative is nobody looking at all. The first useful thing an agent does in most enterprises is not a decision. It is finally reading everything.
The uncomfortable part
An organization that cannot say what a customer is has been running on human judgment it never wrote down and never valued. That judgment was doing real work, invisibly, in thousands of small corrections a day.
The choice is not whether to encode it. The choice is whether to encode it deliberately, now, with the people who still hold it, or to discover it in production as a series of expensive surprises after those people have moved on.
Image The propagation
painterly editorial illustration, a single drop of dark ink spreading outward through the visible fibers of heavy paper, seen close and from above, deep navy ink on warm off white stock, amber lamplight from one side, soft teal in the wet edge, quiet inevitability, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw Your specifics would sharpen this (2)
The piece stands on general enterprise truth. Each line below marks a place where a detail only you have would hit harder. Approximations are fine, labelled as approximations.
- A concrete duplicate-or-drift pattern you have personally untangled, described without naming the employer or the customer.
- Roughly how long a serious master data cleanup took on a project you saw through, as an approximation labelled as one. This number is the article's strongest evidence and I will not invent it.
Art still to generate (3)
Every slot in this piece with no asset yet. Copy a prompt, generate it by hand, commit the file, and its entry disappears from this list.
- One customer, four records Midjourney prompt
painterly editorial illustration, four nearly identical index cards pinned in a row on a dark board, each slightly different in wear and hand, deep navy surfaces, warm amber lamplight, soft teal shadow beneath the cards, tactile paper and brass pins, quiet unease, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw - The readiness test Graph prompt
Draw a clean editorial checklist diagram titled The five question master data readiness test. Five numbered rows, each with a question on the left and a pass condition on the right. Row one, Can you name the system of record for each core entity, pass condition One name, no debate. Row two, Do duplicates have an owner, pass condition A named person, not a committee. Row three, Is there a definition of the entity in writing, pass condition Written, current, and used in disputes. Row four, Can a wrong value be traced to who set it, pass condition Yes, within minutes. Row five, Does anyone measure quality on a schedule, pass condition A recurring number somebody reads. At the bottom a single band reading, Three or more failures means the model is not your constraint. Style, restrained editorial infographic, deep navy and slate on a warm off white ground, one amber accent on the bottom band, thin rules, generous whitespace, sans serif labels, no icons, no gradients, no clutter. - The propagation Midjourney prompt
painterly editorial illustration, a single drop of dark ink spreading outward through the visible fibers of heavy paper, seen close and from above, deep navy ink on warm off white stock, amber lamplight from one side, soft teal in the wet edge, quiet inevitability, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw