Writing

Bringing an Agent Proposal to Finance

The questions I would prepare for before asking someone to sign off

The ledger's memory A ledger and balance scales, the tools of a review that ends with someone signing their name.

A team can bring a working agent to finance and leave the meeting unsure why the proposal was turned down.

The demonstration may have gone well. The team has measured accuracy and investigated the failures it knows about. When finance still says no, it is tempting to look for a more sympathetic sponsor. Before doing that, I would want to understand what the people responsible for the numbers still need to see.

Their questions can change the design, including who approves an output and what evidence the system retains.

Accuracy is part of that discussion. The person signing off also needs to explain how a result was produced and defend it to somebody outside the team. A successful demonstration can leave those requirements unanswered.

Preparing for the review

These are the four questions I would prepare for before bringing an agent proposal to finance.

Can it produce the same answer twice? Tax determination, statutory reporting, revenue recognition, and period close all have to be reproducible. This is not a finance idiosyncrasy: irreproducible results are treated as unreliable wherever the stakes are high enough to check (Kapoor & Narayanan, 2023). The reviewer needs to be able to trace the result to the inputs and rules in force at the time. A good average score does not provide that record.

Two different questions Accuracy, reproducibility, accountability, explanation, and a tested fallback belong in the same design review.

Who is accountable when it is wrong? Every consequential action in a controlled environment answers two questions: who did this, and were they permitted to. The log needs to identify the responsible person and their authority to approve the action. A service account alone leaves that responsibility unclear. Until an agent can stand inside segregation of duties rather than punch a hole through it, its involvement weakens a control that somebody signed for (Committee of Sponsoring Organizations of the Treadway Commission, 2013). What it may reach at all is a least-privilege question (Saltzer & Schroeder, 1975).

Can we explain it to an auditor? The reviewer needs a path from source document to posted number that someone outside the organization can follow. A confidence score and the name of a model leave too much unexplained. Documenting intended use, known limits and evaluation conditions up front is the closest available substitute (Mitchell et al., 2019), and it is broadly what the AI governance frameworks now ask for (National Institute of Standards and Technology, 2023).

Can we turn it off in the middle of a period? Every process that touches close needs a manual fallback that a human can actually execute, tested, with the runbook current. The fallback needs an owner and enough detail for someone to carry out the work during an interruption. For high-risk systems within its scope, Article 14 of the EU AI Act provides for human intervention and interruption of the system (European Parliament & Council, 2024).

Why these are the right questions

The questions make more sense when you consider the work the reviewer is accountable for.

In a ledger that retains posted entries, a correction adds a reversing entry rather than removing the history. That design matters when considering internal control over financial reporting (Public Company Accounting Oversight Board, 2024). I have made this argument at length in the reversibility ladder.

Financial statements are attested. Somebody signs, personally, and the consequences of an inaccurate signature are not a bad quarter. That signature is the reason the questions above sound uncompromising: they are the conditions under which a person is willing to put their name on an output they did not personally produce.

An overall success rate also leaves open the question of where failures occur. Unusual transactions deserve separate attention, especially where several controls depend on the same assumption. The safety literature gives a useful account of how weaknesses in different defences can combine (Reason, 2000).

How to get to yes

Bring a list of proposed operations alongside the accuracy results. Show which operations the agent will touch and what it costs to reverse each one. Stating which operations remain outside the agent's access gives finance a concrete proposal to assess.

Make the agent's output a proposal with its evidence attached. The strongest pattern I know is an agent that produces the recommended entry, the documents it drew on, the rule it applied, and the specific reason, handed to a human who commits it. That gives the reviewer material to check and can reduce the time spent assembling it. The decision and its accountability remain with the person who commits the entry. It also matches what is known about designing for correction: make the evidence visible and the fix cheap (Amershi et al., 2019).

Offer reproducibility deliberately. Version the prompt, the model, and the rule set; store them with the output. When somebody re-runs it in nine months and gets a different answer, you want to be able to say precisely why, and "the vendor updated the model" is a much better answer when you can prove which one you used.

Say what it does not do, first and unprompted. Volunteering the boundary is what buys the trust to expand it later. Programs that oversell their scope in the first meeting spend the following year re-earning credibility they could simply have kept. The asymmetry is measurable: people drop an algorithm sharply after seeing it err, even one that beats them on average (Dietvorst, Simmons & Massey, 2015), and appropriate reliance has to be earned against demonstrated reliability rather than asserted (Lee & See, 2004).

And find the controller who is curious. There is usually one. They have wanted the close to stop consuming their team's evenings for years and they know exactly which parts are assembly rather than judgment. They are also the person who can tell you which of your ideas will die in front of the auditors, before the team spends months building them.

What I would bring back to the team

If those questions are still unanswered, I would bring them back as design work. Some may require a narrower first release. Others may need a record, a review step, or a fallback that has not yet been built.

A capable system can still go unused, a problem the human factors literature describes as disuse (Parasuraman & Riley, 1997). I would rather work through the reviewer's concerns while the scope is still easy to change.

The signature The person signing needs a record they can explain later.

References

  1. Public Company Accounting Oversight Board (2024). AS 2201: An Audit of Internal Control Over Financial Reporting. PCAOB Auditing Standards. pcaobus.org/oversight/standards/auditing-standards/details/AS2201 An auditing standard for internal control over financial reporting. The essay uses it as control context, not as a product-specific ledger specification.
  2. Committee of Sponsoring Organizations of the Treadway Commission (2013). Internal Control, Integrated Framework. COSO. www.coso.org/guidance-on-ic Control environment, segregation of duties, and the monitoring expectations behind these questions.
  3. European Parliament & Council (2024). Article 14: Human oversight, Regulation (EU) 2024/1689 (AI Act). Official Journal of the European Union. eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
  4. National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Department of Commerce. nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
  5. Lee & See (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50-80. doi.org/10.1518/hfes.46.1.50_30392 Calibrated trust: reliance should track demonstrated reliability, not confidence of presentation.
  6. Dietvorst, Simmons & Massey (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1). doi.org/10.1037/xge0000033 People abandon an algorithm faster after seeing it err, even when it outperforms them.
  7. Parasuraman & Riley (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230-253. doi.org/10.1518/001872097778543886
  8. Reason (2000). Human error: models and management. BMJ, 320(7237), 768-770. pmc.ncbi.nlm.nih.gov/articles/PMC1117770
  9. Saltzer & Schroeder (1975). The protection of information in computer systems. Proceedings of the IEEE, 63(9). web.mit.edu/Saltzer/www/publications/protection
  10. Mitchell et al. (2019). Model cards for model reporting. arXiv:1810.03993. arxiv.org/abs/1810.03993 A documentation format that makes intended use and known limits explicit.
  11. Kapoor & Narayanan (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9). doi.org/10.1016/j.patter.2023.100804 Examines reproducibility and leakage problems in machine-learning-based science.
  12. Amershi et al. (2019). Guidelines for Human-AI Interaction. CHI Conference on Human Factors in Computing Systems. dl.acm.org/doi/10.1145/3290605.3300233