Writing
What Finance Actually Needs Before It Will Accept an Agent
The objection is not fear of the technology. It is a different question entirely.
Image The ledger's memory
painterly editorial illustration, a heavy bound accounting ledger closed on a dark desk beside a small set of brass scales, deep navy surfaces, warm amber lamplight raking across leather and paper tooth, soft teal shadow, weight and permanence, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw There is a meeting that happens in every enterprise AI program, and it goes the same way almost everywhere.
The technical team demonstrates something genuinely good. The accuracy is strong, the failure cases are understood, the demo works. Finance says no. The team leaves believing finance is risk averse, behind the times, or protecting territory, and starts looking for a sponsor who will say yes.
That reading is wrong, and it is the most expensive misread in the field.
Finance is not evaluating whether the agent is accurate. It is evaluating whether the agent's output can be defended by a named person to somebody external. Those are different tests, and a system can pass the first perfectly while failing the second completely.
The four questions actually being asked
Underneath a polite "not yet" there are usually four specific questions. None of them is about model quality.
Can it produce the same answer twice? Tax determination, statutory reporting, revenue recognition, and period close all have to be reproducible. Not accurate on average. Reproducible, so that the same inputs six months from now, re-run during an examination, yield the same output with the same reasoning. A probabilistic system that is right more often than the human it replaced is still the wrong shape for that job, because the requirement was never accuracy. It was the ability to re-derive the number on demand.
Figure Two different questions
Draw a clean editorial two column comparison diagram titled Two different questions. Left column headed Engineering asks, with four stacked items, Is it accurate, How often is it wrong, Does it beat the baseline, Can we improve it next release. Right column headed Finance asks, with four stacked items, Can it produce the same answer twice, Who is accountable when it is wrong, Can we explain it to an auditor, Can we turn it off mid period. Draw a vertical rule between the columns. Beneath both, a single full width band reading, A system can pass every item on the left and fail every item on the right. Style, restrained editorial infographic, deep navy and slate on a warm off white ground, one amber accent on the right column heading, thin rules, generous whitespace, sans serif labels, no icons, no gradients, no clutter. Who is accountable when it is wrong? Every consequential action in a controlled environment answers two questions: who did this, and were they permitted to. That is not logging. It is accountability, and it presumes a person who can be asked what they were thinking. A service account running an agent satisfies the field and defeats the purpose. Until an agent can be a party to segregation of duties rather than a hole in it, its involvement weakens a control that somebody signed for.
Can we explain it to an auditor? Not a confidence score. A walkable path from source document to posted number, in language that survives being read by somebody who does not work here and is not obliged to be impressed. "The model determined it" is not an explanation, it is a description of where the explanation should have been.
Can we turn it off in the middle of a period? Every process that touches close needs a manual fallback that a human can actually execute, tested, with the runbook current. Not because the agent is expected to fail. Because the plan for it failing is itself a control, and a process whose only path runs through a system nobody in the room can operate by hand is a single point of failure that happens to be fashionable.
Why these are the right questions
It is tempting to treat this as institutional caution and route around it. Worth understanding why the standard exists before deciding it is obsolete.
A general ledger does not forget. You cannot delete a posting; you post a reversing entry in an open period, and both live in the record permanently. So a mistake is not undone, it is appended, and the correction becomes part of a history somebody will later have to explain. Software people hear "reversible" and think undo. Accounting means something much narrower.
Financial statements are attested. Somebody signs, personally, and the consequences of an inaccurate signature are not a bad quarter. That signature is the reason the questions above sound uncompromising: they are the conditions under which a person is willing to put their name on an output they did not personally produce.
And controls compound. A control that works ninety eight percent of the time is not a ninety eight percent control, because the two percent is not randomly distributed. It concentrates exactly where the process is unusual, which is exactly where the fraud and the errors live. Finance knows this in its bones, from long before software was involved.
How to get to yes
Bring the tiering to the meeting, not the accuracy. Show which operations the agent will touch and what it costs to reverse each one. A team that opens with "it never posts, it never pays, it never closes a period, here is precisely where it stops" has already answered the loudest objection in the room and can spend the rest of the meeting on value.
Make the agent's output a proposal with its evidence attached. The strongest pattern I know is an agent that produces the recommended entry, the documents it drew on, the rule it applied, and the specific reason, handed to a human who commits it. That preserves the control, preserves the accountability, and still removes most of the work, because most of the work was never the decision. It was the assembly.
Offer reproducibility deliberately. Version the prompt, the model, and the rule set; store them with the output. When somebody re-runs it in nine months and gets a different answer, you want to be able to say precisely why, and "the vendor updated the model" is a much better answer when you can prove which one you used.
Say what it does not do, first and unprompted. Volunteering the boundary is what buys the trust to expand it later. Programs that oversell their scope in the first meeting spend the following year re-earning credibility they could simply have kept.
And find the controller who is curious. There is usually one. They have wanted the close to stop consuming their team's evenings for years and they know exactly which parts are assembly rather than judgment. They are also the person who can tell you which of your ideas will die in front of the auditors, which is worth more than any executive sponsor.
The part engineering does not want to hear
If your agent cannot answer those four questions, finance is not the obstacle. The design is incomplete, and finance is the first group to have read it carefully.
The teams that get furthest treat that reading as free design review from the most rigorous stakeholder in the building. The ones that stall treat it as politics, go find a friendlier sponsor, and ship something that quietly gets switched off during the first close that goes badly.
Image The signature
painterly editorial illustration, close study of a fountain pen resting on the signature line of a formal financial filing, deep navy desk, warm amber lamplight, crisp paper with visible tooth, one soft teal reflection on the pen barrel, gravity and consequence, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw Your specifics would sharpen this (2)
The piece stands on general enterprise truth. Each line below marks a place where a detail only you have would hit harder. Approximations are fine, labelled as approximations.
- The clearest example you have of a controller asking a question the technical team had not considered. Anonymized, no employer named.
- Whether you want to name the specific control frameworks you have worked under, or keep this at the level of the underlying principle.
Art still to generate (3)
Every slot in this piece with no asset yet. Copy a prompt, generate it by hand, commit the file, and its entry disappears from this list.
- The ledger's memory Midjourney prompt
painterly editorial illustration, a heavy bound accounting ledger closed on a dark desk beside a small set of brass scales, deep navy surfaces, warm amber lamplight raking across leather and paper tooth, soft teal shadow, weight and permanence, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw - Two different questions Graph prompt
Draw a clean editorial two column comparison diagram titled Two different questions. Left column headed Engineering asks, with four stacked items, Is it accurate, How often is it wrong, Does it beat the baseline, Can we improve it next release. Right column headed Finance asks, with four stacked items, Can it produce the same answer twice, Who is accountable when it is wrong, Can we explain it to an auditor, Can we turn it off mid period. Draw a vertical rule between the columns. Beneath both, a single full width band reading, A system can pass every item on the left and fail every item on the right. Style, restrained editorial infographic, deep navy and slate on a warm off white ground, one amber accent on the right column heading, thin rules, generous whitespace, sans serif labels, no icons, no gradients, no clutter. - The signature Midjourney prompt
painterly editorial illustration, close study of a fountain pen resting on the signature line of a formal financial filing, deep navy desk, warm amber lamplight, crisp paper with visible tooth, one soft teal reflection on the pen barrel, gravity and consequence, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw