Writing
Before I Let an Agent Do More
How I would examine the tests behind a proposal to expand its authority
Before agreeing to give an agent access to a system of record, I would ask the team a question.
How do you know it works?
A demo and a period of informal use can give a team confidence, but I would want to see how that confidence was tested. Occasionally there is a number, and the number came from a set of cases the same team assembled and graded, chosen partly because they were the easiest cases to collect. The literature on benchmarking has a name for how much that choice can decide: the ranking often follows the benchmark rather than the method (Dehghani et al., 2021).
An evaluation gives the process owner something to examine before extending the agent's authority. It should show the cases tested, the answers expected, the mistakes made, and what those mistakes would cost.
Why a demo cannot answer the question
A demonstration shows selected cases, usually chosen by the team that built the system. It answers "can this ever work," which is a real question and one you should ask early.
It cannot answer any of the questions that determine whether the thing should be allowed near production: how often it is wrong, what kind of wrong, whether the wrongness clusters somewhere consequential, and whether any of that changed since last month. Aggregate scores hide exactly this structure, which is why holistic and behavioral evaluation exist as separate disciplines (Liang et al., 2022) (Ribeiro et al., 2020).
The result can also change after the evaluation. Everything underneath the system moves. Vendors can update models without a version change you notice. Prompts get edited by whoever was closest to the complaint. The data drifts, new product lines appear, a business rule changes. An earlier measurement may no longer describe the current system, even if the team still relies on it. Contamination and leakage make the decay worse, because a stale set quietly stops testing what it was built to test (Kapoor & Narayanan, 2022).
The four questions
What is the correct answer, and who decided? The people accountable for the process need a say in what counts as correct. If the build team sets and grades every case alone, the evaluation can reproduce assumptions that no one has independently examined. General-purpose benchmarks tend to claim more generality than their construction can support, and enterprise-built ones inherit the same flaw in miniature (Raji et al., 2021).
In enterprise settings the answer usually already exists in the organization: the controller, the master data owner, the planner who has done this for eleven years. Their judgment is a starting point for agreed answers, and getting time with them belongs in the project plan. It is also where you discover that two experts disagree on a surprising share of the cases, which is not a problem with your eval. Disagreement among qualified raters is a measurable property of the task, and it belongs in the report rather than being averaged away (Zheng et al., 2023). Those disagreements can reveal decisions that the process documentation leaves open.
Where did the cases come from? A set assembled from whatever was convenient will overrepresent the ordinary. Make room for the awkward cases: the return that spans a period boundary, the customer who is also a supplier, the item with a unit of measure conversion, the intercompany case, the country with the local rule. If your set has no cases that make an expert pause, it is measuring the part that was never in doubt. How a set was assembled is itself evidence, and worth documenting as such (Gebru et al., 2021).
What does wrong cost here? Counting errors equally can hide a costly failure among many harmless ones. A missed classification that a human catches in review is not the same event as a confidently wrong value that propagates into a planning run. Weight by consequence, or you will optimize a number that improves while the risk gets worse. Broad multi-task scores are particularly prone to this, since they average very different kinds of failure into one figure (Hendrycks et al., 2020).
Would it still pass next month? Which means it has to be rerunnable, on a schedule, by someone who was not present at the build. This is an engineering property more than an analytical one, and it is the difference between an eval and a study. Production readiness rubrics treat rerunnable tests as a requirement rather than a nicety (Breck et al., 2017).
What good looks like in practice
I would start with a manageable set of difficult cases and an owner outside the build team. A set of well argued cases with expert-agreed answers will tell you more than a large set assembled without argument, and it can exist in a week rather than a quarter.
Versioned alongside the thing it measures. The prompt, the model, the rules, and the case set are one artifact with one version. Recording intended use and evaluation conditions alongside the artifact is now standard practice (Mitchell et al., 2019), and governance frameworks expect the measurement to be maintained rather than performed once (National Institute of Standards and Technology, 2023). When a score moves, the first question is which of the four changed, and that question needs an answer that is a lookup rather than an investigation.
Run automatically and reported to somebody who is not on the team. The process owner needs the results and a way to act on them. Decide who reviews a changed score and what would cause them to stop or narrow the release.
Complete with a written disagreement log. When two experts differed, record it and record how it was resolved. That record can help a later reviewer understand a decision without having to find everyone who attended the original discussion.
When experts disagree
For a large share of interesting enterprise work, there is no clean ground truth. The right answer to "should this order ship partial" or "is this variance worth investigating" depends on context, on relationships, and on what the business is optimizing this quarter.
Benchmarking itself has unresolved structural problems (Bowman & Dahl, 2021). What can be done is narrower and still worth a great deal: measure the parts that are decidable, measure the disagreement on the parts that are not, and be explicit about which is which. A report that says "these forty cases have agreed answers and we score them, these fifteen are contested and here is the spread" is more credible, and more useful, than one reporting a single confident number across both.
Alongside its use in tuning, the evaluation helps the accountable person decide whether to widen the agent's authority. They need to see which decisions have agreed answers and which still require judgment.
Before the next release
I would use the same questions when assessing a proposal from another team and when preparing one of my own.
Ask to see the case set, who agreed the answers, how errors are weighted, and when the tests last ran. A score is easier to interpret once those details are available.
If those details are missing, I would spend time establishing them before proposing a larger use case. The result might support expansion, or it might show us what to fix first.
References
- Liang et al. (2022). Holistic Evaluation of Language Models. arXiv:2211.09110. arxiv.org/abs/2211.09110
- Raji et al. (2021). AI and the everything in the whole wide world benchmark. arXiv:2111.15366. arxiv.org/abs/2111.15366 Argues that general benchmarks routinely claim more generality than their construction supports.
- Bowman & Dahl (2021). What will it take to fix benchmarking in natural language understanding?. arXiv:2104.02145. arxiv.org/abs/2104.02145 Discusses limitations in benchmark design and the interpretation of reported scores.
- Dehghani et al. (2021). The benchmark lottery. arXiv:2107.07002. arxiv.org/abs/2107.07002 Examines how benchmark and configuration choices can affect comparisons between methods.
- Zheng et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv:2306.05685. arxiv.org/abs/2306.05685 Using a model as judge, including its measured agreement with human raters and its biases.
- Ribeiro et al. (2020). Beyond accuracy: Behavioral testing of NLP models with CheckList. arXiv:2005.04118. arxiv.org/abs/2005.04118 Behavioral testing: capability-specific cases rather than a single aggregate score.
- Kapoor & Narayanan (2022). Leakage and the reproducibility crisis in ML-based science. arXiv:2207.07048. arxiv.org/abs/2207.07048
- Mitchell et al. (2019). Model cards for model reporting. arXiv:1810.03993. arxiv.org/abs/1810.03993
- Gebru et al. (2021). Datasheets for datasets. arXiv:1803.09010. arxiv.org/abs/1803.09010 Proposes documenting dataset collection, composition, and intended uses.
- Hendrycks et al. (2020). Measuring massive multitask language understanding. arXiv:2009.03300. arxiv.org/abs/2009.03300
- Breck et al. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. IEEE International Conference on Big Data. doi.org/10.1109/BigData.2017.8258038
- National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Department of Commerce. nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf