Writing
Before an Agent Reads the Customer Records
Finding the corrections people make before trusting the data
An AI readiness review needs to look beyond infrastructure and skills. I would also ask the team to follow a customer record through the systems the proposed agent will use.
That exercise can expose disagreement about what counts as one customer. Data quality has never been a single number: it is multidimensional and defined by fitness for use, which is why two teams can both be right about the same field (Wang & Strong, 1996).
Master data deserves an owner and a place in the readiness review. You can examine its quality before choosing a model, including the exceptions people currently resolve for themselves.
Why this bites agents harder than it bit reports
Enterprises have lived with imperfect master data for decades. It is survivable because humans absorb it silently.
A person pulling a report sees three rows for what is obviously one supplier and mentally merges them. A person entering an order recognizes that the customer with the trailing period in the name is the same account as the one without. A planner knows that the item flagged as active has not been purchased in four years and skips it. Some of that judgment may never have made it into documentation, even though people rely on it throughout the day. This is tacit knowledge in the ordinary organizational sense, and its defining property is that it resists being written down at all (Nonaka, 1994).
An agent cannot be assumed to make the same corrections. If the workflow depends on a planner recognizing an obsolete item, we need to decide how that check will happen when the planner is no longer reading each row.
A reporting process can conceal that dependence for years. Removing the informal checks exposes errors that people previously corrected. The operational costs of poor data quality have been documented for decades (Redman, 1998), and in machine learning specifically it compounds into cascades that surface far downstream (Sambasivan et al., 2021).
The five question test
These questions can give an initial view of the work required. Answer them with the people who maintain the records.
Can you name the system of record for each core entity? Customer, supplier, item, employee, chart of accounts. One name each, said quickly, without a debate breaking out. If three systems create customers and reconciliation runs nightly between them, ask which record governs each decision and how conflicts are resolved.
Do duplicates have an owner? Name a person whose job includes the decision, with a clear route for resolving disputes. Large enterprises reliably have duplicate master records. The healthy ones have somebody whose week includes merging them, with the authority to decide which survives.
Is the entity defined in writing? What makes a customer one customer. A legal entity, a ship-to, a billing relationship, a parent group. Four teams can answer that question differently for valid local reasons. Those differences need to be explicit before their records are treated as interchangeable.
Can a wrong value be traced to who set it? Given a bad credit limit or a wrong tax group, can somebody find out who changed it, when, and from what, within minutes. If the answer is a database restore, an agent that touches this data cannot be governed, because you will be unable to distinguish its mistakes from anyone else's. Traceability of who changed what is a control expectation, not a nicety (Committee of Sponsoring Organizations of the Treadway Commission, 2013).
Does anybody measure quality on a schedule? A recurring number that a named person actually reads. Duplicate rate, completeness on required fields, records untouched for a suspicious length of time. Without a recurring measure, new errors can accumulate unnoticed while attention stays on creating records. The empirical work on inventory records is the clearest illustration available: records and physical reality diverge continuously, at scale, in well-run operations (DeHoratius & Raman, 2008). Data validation belongs in the production readiness checklist for exactly this reason (Breck et al., 2017).
Several unresolved answers would make me pause before granting write access. Deployment studies keep finding the same thing: the binding problems sit in data and process rather than in modelling (Paleyes, Urma & Lawrence, 2022). The required cleanup may take several quarters and need its own owner and budget.
Why the cleanup keeps getting deferred
People may already know where the records are unreliable. Getting the work funded requires explaining why the existing incentives have left it undone.
The cost is borne centrally and the benefit lands everywhere else. If the benefit does not appear in the responsible team's objectives, other work can keep taking priority.
It has no end state, which makes it hostile to project funding. Master data quality is a maintained property, like cleanliness. Data dependencies behave like debt, accruing interest quietly until something downstream breaks (Sculley et al., 2015). Programs that treat it as a one-time cleanup deliver a genuine improvement and then watch it decay, which teaches the organization that the work does not stick, which makes the next cleanup harder to fund.
And it is politically expensive in a way technical work usually is not. Deciding what a customer is means telling several teams that their local definition, which works fine for them, is no longer the one that counts. The decision needs someone with authority across the affected teams.
What to do instead of a boil-the-ocean program
Scope the cleanup to the agent's actual reach. You do not need enterprise-wide perfection. You need the entities that this use case reads and writes to be trustworthy, which is a far smaller and far more fundable problem. Let the first project pay for the first slice.
The process that creates duplicates matters more than the backlog of them. Cleaning duplicates while the process that generates them runs untouched is a treadmill. Validation at entry, a required search-before-create, and a single owning path are less satisfying than a big remediation number and worth considerably more.
Before remediating anything, publish the duplicate rate and the completeness numbers. Governance frameworks put measurement ahead of management for the same reason (National Institute of Standards and Technology, 2023). It gives you a baseline, it makes the improvement visible to people who fund things, and it is the only way to notice the decay later.
An agent can also help identify records that need review. Reconciliation, duplicate detection, and completeness checking are Tier 1 work: nothing is committed, a wrong answer costs somebody's attention, and a person can review the suggested correction before it changes a record. Reviewing large volumes of records can be a useful early application as model capability and cost change (Stanford Institute for Human-Centered AI, 2026).
Begin with the people maintaining the records
When teams use different customer definitions, the people reconciling their work may already know how to resolve many of the differences. Give that knowledge a place in the cleanup plan.
I would begin with the records the first agent will touch and the people who correct them now. Ask what they check, record the exceptions, and give someone time to maintain those decisions. That is a manageable place to start even when the wider cleanup will take much longer.
References
- Sambasivan et al. (2021). Everyone wants to do the model work, not the data work: Data cascades in high-stakes AI. CHI Conference on Human Factors in Computing Systems. doi.org/10.1145/3411764.3445518 Data cascades: upstream data problems surface late, downstream, and expensively.
- Wang & Strong (1996). Beyond accuracy: What data quality means to data consumers. Journal of Management Information Systems, 12(4). doi.org/10.1080/07421222.1996.11518099 Establishes data quality as multidimensional and defined by fitness for use, not accuracy alone.
- Redman (1998). The impact of poor data quality on the typical enterprise. Communications of the ACM, 41(2), 79-82. doi.org/10.1145/269012.269025 Early accounting of how poor data quality propagates into operational and strategic cost.
- DeHoratius & Raman (2008). Inventory record inaccuracy: An empirical analysis. Management Science, 54(4), 627-641. doi.org/10.1287/mnsc.1070.0789 Empirical study of inventory record inaccuracy across a retailer's stores.
- Sculley et al. (2015). Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28. proceedings.neurips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html
- Paleyes, Urma & Lawrence (2022). Challenges in Deploying Machine Learning: A Survey of Case Studies. ACM Computing Surveys, 55(6). arxiv.org/abs/2011.09926
- Nonaka (1994). A dynamic theory of organizational knowledge creation. Organization Science, 5(1), 14-37. doi.org/10.1287/orsc.5.1.14 Discusses tacit knowledge and its relationship to organizational knowledge creation.
- Committee of Sponsoring Organizations of the Treadway Commission (2013). Internal Control, Integrated Framework. COSO. www.coso.org/guidance-on-ic
- National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Department of Commerce. nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
- Breck et al. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. IEEE International Conference on Big Data. doi.org/10.1109/BigData.2017.8258038
- Stanford Institute for Human-Centered AI (2026). AI Index Report. Stanford University. hai.stanford.edu/ai-index