Writing
Getting an AI Pilot Into Production
The integration work I would include in the first estimate
When an enterprise AI pilot stalls, model accuracy is one place to look. I would also examine the work needed to connect it to the systems people use each day.
I have watched enough pilots encounter those difficulties that I would include them in the first estimate, while the scope is still open.
A useful demonstration can leave identity, data contracts, support, and exception handling largely untested. The next stage needs time and people for that work.
What the demonstration depended on
A pilot gets built quickly, and it gets built quickly for a reason. The team may have exported a spreadsheet, selected the records, and run against a fixed copy of the data using a developer's own account.
Those shortcuts can be useful for testing an idea. Record them so the next estimate includes replacing the ones production cannot use.
A successful demonstration can make the remaining work look smaller than it is. Reviews of deployments describe substantial difficulties beyond the model: the obstacles cluster in data handling, infrastructure and organizational fit rather than in modelling (Paleyes, Urma & Lawrence, 2022).
The integration estimate needs to account for those dependencies before the team promises a production date.
Work to include in the estimate
These are seven areas I would discuss with the system owners before expanding the pilot.
Identity. The agent needs an account with defined access, following the principle of least privilege (Saltzer & Schroeder, 1975). The access reviewer needs to approve the permissions and know who owns the account. In a controlled environment the question is not how to create a service account. It is who approves it, what it may reach, whether it can be told apart from a person in the audit trail, and what happens at the next access recertification when its owner has moved teams.
Data contracts. The demo read a field. Production has four fields that look like that one, populated by three different processes, two of which are load-bearing and one of which is a convention nobody wrote down. I have seen those differences expand the scope of an integration. The research describes a related problem in data cascades, whose damage surfaces late and far downstream (Sambasivan et al., 2021).
Environments and test data. To test an integration you need a place to break it. In enterprise systems that place is a lower environment whose data is a partial, aging, differently-shaped copy of production, and whose refresh cadence is owned by somebody with other priorities. In my experience a surprising amount of program time is spent not building and not testing, but waiting for somewhere legitimate to try.
Error handling. A demo has one path. An interface has a failure taxonomy: the record was locked, the period was closed, the value was rejected downstream, the call timed out, the call succeeded and the response was lost. Each one needs a decision about retry, alert, and whether a partial write is worse than no write. A production readiness rubric can help the team work through these cases before promising a date (Breck et al., 2017).
Monitoring and ownership. The pilot team goes back to their normal jobs. Something has to page somebody when the run fails at two in the morning, and that somebody has to have a runbook they did not write. Agree who will support it after the pilot team moves on, including who has time to maintain the runbook. The ongoing burden is real and it accrues quietly: maintenance debt in machine learning systems is dominated by everything surrounding the model (Sculley et al., 2015).
Windows. Batch schedules, close periods, freeze windows, statutory cutoffs, regional cutovers. In my experience a system that is technically available is not always organizationally open, and an agent with no model of the calendar will eventually act inside a window where every human involved has deliberately stopped.
Change management. The people whose work the agent touches have to change what they do. That is its own discipline and its own article (Kotter, 1995), and it can determine whether a technically successful integration changes the work.
Why the framing keeps repeating
The integration work can be hard to explain during a funding discussion.
A model result is easier to demonstrate than weeks of integration work. You can show a stakeholder a summarized variance in forty seconds; you cannot show them the six weeks of reconciliation logic that will make the summary trustworthy on the fourth Tuesday of a close. So the budget follows the demonstrable thing, the timeline is set by the demonstrable thing, and the rest arrives as a surprise that looks like failure.
Staffing also matters. Integration work is the part of the program that requires knowing the specific system, and specific system knowledge is exactly what an AI initiative staffed as an AI initiative tends not to have on it. The resulting architecture tends to mirror the org chart that produced it rather than the problem (Conway, 1968), and the difficulty is essential rather than accidental, so no tool removes it (Brooks, 1987). The people who know why that field has four variants are usually in a different part of the organization, doing something else, and were not invited.
Planning the first production release
For a first release, I would favor a useful task with a small, well-understood integration. That gives the team a chance to test the whole operating process, from access and data quality through correction and support.
Budget the seven things above explicitly, as line items, before the pilot starts. The governance frameworks are useful here precisely because they force the mapping to be written down (National Institute of Standards and Technology, 2023). Leave the unresolved items visible in the estimate, with someone responsible for following them up.
Involve the people who maintain the source systems while choosing the scope. They can explain which field to use, what its exceptions mean, and who else depends on it.
Use the smaller first release to establish the supporting work. Logging, replay, audit, alerting, a place to test. Design the correction path while you are at it, because a system people cannot efficiently correct is one they quietly stop using (Amershi et al., 2019). You will need all of it for the valuable use case later, and you would much rather build it while the stakes are low and nobody is watching.
And say out loud, early, that the first release will look modest. Explain what the first release will establish and what remains outside its scope.
What the team can work on now
An AI proposal may depend on integration, data quality, or ownership problems left unresolved by earlier projects. Naming them helps the team decide whether the proposed release is feasible.
Model capability continues to change on vendors' release schedules (Stanford Institute for Human-Centered AI, 2026). Tight coupling between the systems you are joining, meanwhile, is a property you own (Perrow, 1999). I would put the integration dependencies beside the model requirements and plan the release with the people who can address both.
References
- Paleyes, Urma & Lawrence (2022). Challenges in Deploying Machine Learning: A Survey of Case Studies. ACM Computing Surveys, 55(6). arxiv.org/abs/2011.09926 Survey of deployment case studies; the reported obstacles cluster outside model quality.
- Sculley et al. (2015). Hidden Technical Debt in Machine Learning Systems. Advances in Neural Information Processing Systems 28. proceedings.neurips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html Names the ongoing costs a prototype defers: glue code, pipeline jungles, configuration debt.
- Breck et al. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. IEEE International Conference on Big Data. doi.org/10.1109/BigData.2017.8258038 A production readiness rubric for machine learning systems, including testing and monitoring.
- Conway (1968). How do committees invent?. Datamation, April 1968. www.melconway.com/Home/Committees_Paper.html Systems mirror the communication structure of the organizations that build them.
- Brooks (1987). No Silver Bullet: Essence and accidents of software engineering. Computer, 20(4), 10-19. doi.org/10.1109/MC.1987.1663532 The essential-versus-accidental distinction this article applies to the integration layer.
- Amershi et al. (2019). Guidelines for Human-AI Interaction. CHI Conference on Human Factors in Computing Systems. dl.acm.org/doi/10.1145/3290605.3300233
- National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Department of Commerce. nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
- Stanford Institute for Human-Centered AI (2026). AI Index Report. Stanford University. hai.stanford.edu/ai-index
- Sambasivan et al. (2021). Everyone wants to do the model work, not the data work: Data cascades in high-stakes AI. CHI Conference on Human Factors in Computing Systems. doi.org/10.1145/3411764.3445518
- Perrow (1999). Normal Accidents: Living with High-Risk Technologies. Princeton University Press. press.princeton.edu/books/paperback/9780691004129/normal-accidents
- Saltzer & Schroeder (1975). The protection of information in computer systems. Proceedings of the IEEE, 63(9). web.mit.edu/Saltzer/www/publications/protection Sets out the principle of least privilege used here when discussing access boundaries.
- Kotter (1995). Leading change: Why transformation efforts fail. Harvard Business Review, 73(2). hbr.org/1995/05/leading-change-why-transformation-efforts-fail-2 Discusses organizational change, including short-term wins and the work needed to sustain them.