Writing

Pilots Die at Integration, Not Intelligence

The demo ran on a spreadsheet. Production does not have one.

Image The seam

Midjourney prompt
painterly editorial illustration, two vast industrial machines separated by a gap, joined only by one narrow bridge of bundled cabling, deep navy surfaces, warm amber light from inside the machines, soft teal shadow in the gap, tactile metal and paper, quiet confident order, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw
The seam The distance between a working prototype and a production system is measured in interfaces, not in model quality.

Ask why an enterprise AI pilot failed and you will usually be told a story about the model. It hallucinated. It could not handle the edge cases. The accuracy was not there yet. Wait for the next release.

I have watched enough of these to think that story is mostly wrong, and expensively so, because it points the next attempt at the same wall.

Pilots rarely die of insufficient intelligence. They die at the seam, and the seam is almost never in the pilot's budget.

What the demo was actually standing on

A pilot gets built quickly, and it gets built quickly for a reason. Somebody exported data to a spreadsheet or a flat file. Somebody hand-picked the records. Somebody ran it against a copy that stopped changing three weeks ago. Somebody authenticated as themselves.

None of that is dishonest. It is how you find out whether the idea has legs, and it is the correct way to spend the first two weeks.

The problem is what happens next. The demo goes well. Everyone in the room concludes that the hard part is done, because the visible part is done. And the visible part was the model.

Then the program meets the interface layer, and discovers that the thing it demonstrated in two weeks needs eleven more to exist.

The seven things that were not in the budget

Every one of these is unglamorous, every one is real, and every one is discovered rather than planned in a program that budgeted for intelligence.

Figure Where the work actually is

Graph prompt
Draw a clean editorial diagram titled Where an enterprise AI program actually spends itself. Show one small labeled block at the top called Model and prompt work. Beneath it show a much larger stack of seven labeled blocks of roughly equal size, in this order, Identity and permissions, Data contracts and master data, Environment and test data, Error handling and retries, Monitoring and support ownership, Batch windows and scheduling, Change management and training. Use a bracket down the left side spanning the seven lower blocks labeled Not usually in the pilot budget. Style, restrained editorial infographic, deep navy and slate on a warm off white ground, one amber accent on the small top block, thin rules, generous whitespace, sans serif labels, no icons, no gradients, no clutter.
Where the work actually is The model is the small block. Everything around it is the program.

Identity. The agent needs to be somebody. Not metaphorically. It needs an account, that account needs permissions, and those permissions have to survive a conversation with whoever owns access review. In a controlled environment the question is not how to create a service account. It is who approves it, what it may reach, whether it can be told apart from a person in the audit trail, and what happens at the next access recertification when its owner has moved teams.

Data contracts. The demo read a field. Production has four fields that look like that one, populated by three different processes, two of which are load-bearing and one of which is a convention nobody wrote down. This is the single most reliable source of scope expansion I have seen, and it is not a data science problem. It is a master data problem wearing a data science costume.

Environments and test data. To test an integration you need a place to break it. In enterprise systems that place is a lower environment whose data is a partial, aging, differently-shaped copy of production, and whose refresh cadence is owned by somebody with other priorities. A surprising amount of program time is spent not building and not testing, but waiting for somewhere legitimate to try.

Error handling. A demo has one path. An interface has a failure taxonomy: the record was locked, the period was closed, the value was rejected downstream, the call timed out, the call succeeded and the response was lost. Each one needs a decision about retry, alert, and whether a partial write is worse than no write. This is most of the real engineering, and none of it is interesting to watch.

Monitoring and ownership. The pilot team goes back to their normal jobs. Something has to page somebody when the run fails at two in the morning, and that somebody has to have a runbook they did not write. Programs that cannot answer "who supports this in a year" are not shipping a capability, they are shipping a future outage with a champion attached.

Windows. Batch schedules, close periods, freeze windows, statutory cutoffs, regional cutovers. A system that is technically available is not always organizationally open, and an agent with no model of the calendar will eventually act inside a window where every human involved has deliberately stopped.

Change management. The people whose work the agent touches have to change what they do. That is its own discipline and its own article, and it is the most common reason a technically successful integration produces no measurable value.

Why the framing keeps repeating

The seam is invisible in exactly the moments when programs are funded.

A model is demonstrable in a meeting. An interface is not. You can show a stakeholder a summarized variance in forty seconds; you cannot show them the six weeks of reconciliation logic that will make the summary trustworthy on the fourth Tuesday of a close. So the budget follows the demonstrable thing, the timeline is set by the demonstrable thing, and the rest arrives as a surprise that looks like failure.

There is a second reason, and it is less comfortable. Integration work is the part of the program that requires knowing the specific system, and specific system knowledge is exactly what an AI initiative staffed as an AI initiative tends not to have on it. The people who know why that field has four variants are usually in a different part of the organization, doing something else, and were not invited.

What a program that survives does differently

Pick the use case for its seam, not for its sizzle. The right first project is one where the integration surface is small and well understood, even if the intelligence involved is unimpressive. You are not proving the model works. Everyone already believes the model works. You are proving your organization can run one of these in production.

Budget the seven things above explicitly, as line items, before the pilot starts. A plan that does not name identity, environments, and support ownership has not been costed; it has been hoped at.

Get the specific-system people in the room in week one. Not as reviewers at the end. The person who can tell you which of the four fields is the real one will save more program time in an afternoon than a model upgrade will save in a quarter.

Build the boring substrate on the cheap use case. Logging, replay, audit, alerting, a place to test. You will need all of it for the valuable use case later, and you would much rather build it while the stakes are low and nobody is watching.

And say out loud, early, that the first release will look modest. A program that promises transformation in the first quarter has spent the credibility it will need in the third.

The part nobody wants in the steering deck

The uncomfortable read is that most enterprise AI programs are not blocked on AI at all. They are blocked on the same integration, data quality, and organizational ownership problems that blocked the last three initiatives, which are unfashionable and were never solved.

That is bad news for the roadmap and good news for the odds. Model capability is somebody else's release schedule and you cannot influence it. The seam is entirely yours, and it has been sitting there the whole time, waiting for somebody to treat it as the actual work.

Image The second system

Midjourney prompt
painterly editorial illustration, a quiet operations desk at night, a monitoring board showing one failed run among many successful ones, deep navy room, warm amber desk lamp, one small teal indicator, cold coffee and paper runbooks, restrained and unglamorous, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw
The second system Every enterprise integration eventually becomes something a team maintains at two in the morning.
Your specifics would sharpen this (2)

The piece stands on general enterprise truth. Each line below marks a place where a detail only you have would hit harder. Approximations are fine, labelled as approximations.

  • Roughly how many separate integration surfaces did a typical D365 instance you worked on carry, counting inbound interfaces, outbound feeds, and middleware hops? An approximation is fine, labelled as one.
  • A moment where a working prototype met the real interface layer and the estimate changed. Told without naming the employer or anyone in it.
Art still to generate (3)

Every slot in this piece with no asset yet. Copy a prompt, generate it by hand, commit the file, and its entry disappears from this list.

  1. The seam
    Midjourney prompt
    painterly editorial illustration, two vast industrial machines separated by a gap, joined only by one narrow bridge of bundled cabling, deep navy surfaces, warm amber light from inside the machines, soft teal shadow in the gap, tactile metal and paper, quiet confident order, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw
  2. Where the work actually is
    Graph prompt
    Draw a clean editorial diagram titled Where an enterprise AI program actually spends itself. Show one small labeled block at the top called Model and prompt work. Beneath it show a much larger stack of seven labeled blocks of roughly equal size, in this order, Identity and permissions, Data contracts and master data, Environment and test data, Error handling and retries, Monitoring and support ownership, Batch windows and scheduling, Change management and training. Use a bracket down the left side spanning the seven lower blocks labeled Not usually in the pilot budget. Style, restrained editorial infographic, deep navy and slate on a warm off white ground, one amber accent on the small top block, thin rules, generous whitespace, sans serif labels, no icons, no gradients, no clutter.
  3. The second system
    Midjourney prompt
    painterly editorial illustration, a quiet operations desk at night, a monitoring board showing one failed run among many successful ones, deep navy room, warm amber desk lamp, one small teal indicator, cold coffee and paper runbooks, restrained and unglamorous, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw