Writing
Evals Are the Interview Question You Will Fail
Everyone says they evaluate. Almost nobody can describe what they measured.
Image The gauge
painterly editorial illustration, a single precision pressure gauge mounted on an aged industrial panel, needle resting mid dial, deep navy metal, warm amber lamplight on the brass bezel, soft teal reflection in the glass, tactile rivets and worn paint, quiet exactitude, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw Here is a question I would ask any team proposing to put an agent into a system of record, and it is not a trick.
How do you know it works?
The answer is usually some combination of a demo, a period of informal use, and a genuine, sincerely held feeling that it is good. Occasionally there is a number, and the number came from a set of cases the same team assembled and graded, chosen partly because they were the easiest cases to collect.
Evaluation is the discipline that separates a program that can expand from one that stalls after its first release, and most enterprise AI work has none.
Why a demo cannot answer the question
A demo is a sample of size one, drawn by an interested party, from a distribution they control. It answers "can this ever work," which is a real question and one you should ask early.
It cannot answer any of the questions that determine whether the thing should be allowed near production: how often it is wrong, what kind of wrong, whether the wrongness clusters somewhere consequential, and whether any of that changed since last month.
Figure Four questions an eval must answer
Draw a clean editorial table diagram titled Four questions an evaluation has to answer. Four numbered rows, each with the question on the left and the common failure on the right. Row one, question What is the correct answer and who decided, failure The team grading is the team building. Row two, question Where did the cases come from, failure Cases were chosen because they were easy to collect. Row three, question What does wrong cost here, failure All errors counted equally regardless of consequence. Row four, question Would this still pass next month, failure No rerun, so drift is invisible until a user finds it. Style, restrained editorial infographic, deep navy and slate on a warm off white ground, one amber accent on the failure column heading, thin rules, generous whitespace, sans serif labels, no icons, no gradients, no clutter. That last one is the underrated failure. Everything underneath the system moves. Vendors update models, sometimes without a version change you notice. Prompts get edited by whoever was closest to the complaint. The data drifts, new product lines appear, a business rule changes. A system that was measured once was measured against a world that no longer exists, and the measurement decays silently while confidence in it does not.
The four questions
What is the correct answer, and who decided? This is the hard one and it is where most attempts quietly fail. Somebody has to define correctness, and it cannot be the people who built the thing, because they will encode their own assumptions on both sides of the test and get a very good score.
In enterprise settings the answer usually already exists in the organization: the controller, the master data owner, the planner who has done this for eleven years. Their judgment is the ground truth, and getting a few hours of it is the highest value activity in the whole practice. It is also where you discover that two experts disagree on a fifth of the cases, which is not a problem with your eval. It is the most important thing you learned all quarter, because it means the process you are automating was never as deterministic as its documentation implied.
Where did the cases come from? A set assembled from whatever was convenient will overrepresent the ordinary. The value is in the awkward: the return that spans a period boundary, the customer who is also a supplier, the item with a unit of measure conversion, the intercompany case, the country with the local rule. If your set has no cases that make an expert pause, it is measuring the part that was never in doubt.
What does wrong cost here? Counting all errors equally is the most common analytical mistake in the field. A missed classification that a human catches in review is not the same event as a confidently wrong value that propagates into a planning run. Weight by consequence, or you will optimize a number that improves while the risk gets worse.
Would it still pass next month? Which means it has to be rerunnable, on a schedule, by someone who was not present at the build. This is an engineering property more than an analytical one, and it is the difference between an eval and a study.
What good looks like in practice
Small, hard, and owned by somebody outside the build team. A set of well argued cases with expert-agreed answers will tell you more than a large set assembled without argument, and it can exist in a week rather than a quarter.
Versioned alongside the thing it measures. The prompt, the model, the rules, and the case set are one artifact with one version. When a score moves, the first question is which of the four changed, and that question needs an answer that is a lookup rather than an investigation.
Run automatically and reported to somebody who is not on the team. A score nobody outside the build sees is a private encouragement. The moment it goes to the process owner it becomes a control, and the practice survives the next reorganization.
Complete with a written disagreement log. When two experts differed, record it and record how it was resolved. That log becomes the most valuable document in the program, because it is the only written record of judgment that previously existed only in people, and it will outlive the specific model it was built to test.
The part that is genuinely hard, stated plainly
For a large share of interesting enterprise work, there is no clean ground truth. The right answer to "should this order ship partial" or "is this variance worth investigating" depends on context, on relationships, and on what the business is optimizing this quarter.
Anyone who tells you that is a solved problem is selling. What can be done is narrower and still worth a great deal: measure the parts that are decidable, measure the disagreement on the parts that are not, and be explicit about which is which. A practice that honestly reports "these forty cases have agreed answers and we score them, these fifteen are contested and here is the spread" is more credible, and more useful, than one reporting a single confident number across both.
The credibility matters, because the eval's real function inside an enterprise is not tuning. It is the artifact that lets somebody with accountability decide whether to widen the agent's authority. Without it that decision has no basis but enthusiasm, which is why so many programs get one release and then sit.
Why I called it an interview question
Because it is one, in both directions.
A team that can describe its case set, its ground truth source, its weighting, and its rerun cadence is telling you something real about how it works, regardless of the score. A team whose answer is a demo and a feeling has told you something real as well.
And the same question is worth asking of yourself before proposing the next expansion of scope. If the honest answer is that you do not know how well it works, the correct next project is not a bigger use case.
Image The reference set
painterly editorial illustration, close study of a small set of polished brass calibration weights seated in a fitted velvet lined case, deep navy velvet, warm amber lamplight on brass, soft teal shadow in the recesses, precision and care, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw Your specifics would sharpen this (2)
The piece stands on general enterprise truth. Each line below marks a place where a detail only you have would hit harder. Approximations are fine, labelled as approximations.
- This piece describes what a defensible eval practice looks like, not results from one you have run. If you want a companion piece with your own numbers later, that one has to wait for the numbers.
- Whether to include the hiring frame at all. It is a strong hook for a director audience and it also invites the question of your own eval work, which is currently in progress rather than finished.
Art still to generate (3)
Every slot in this piece with no asset yet. Copy a prompt, generate it by hand, commit the file, and its entry disappears from this list.
- The gauge Midjourney prompt
painterly editorial illustration, a single precision pressure gauge mounted on an aged industrial panel, needle resting mid dial, deep navy metal, warm amber lamplight on the brass bezel, soft teal reflection in the glass, tactile rivets and worn paint, quiet exactitude, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw - Four questions an eval must answer Graph prompt
Draw a clean editorial table diagram titled Four questions an evaluation has to answer. Four numbered rows, each with the question on the left and the common failure on the right. Row one, question What is the correct answer and who decided, failure The team grading is the team building. Row two, question Where did the cases come from, failure Cases were chosen because they were easy to collect. Row three, question What does wrong cost here, failure All errors counted equally regardless of consequence. Row four, question Would this still pass next month, failure No rerun, so drift is invisible until a user finds it. Style, restrained editorial infographic, deep navy and slate on a warm off white ground, one amber accent on the failure column heading, thin rules, generous whitespace, sans serif labels, no icons, no gradients, no clutter. - The reference set Midjourney prompt
painterly editorial illustration, close study of a small set of polished brass calibration weights seated in a fitted velvet lined case, deep navy velvet, warm amber lamplight on brass, soft teal shadow in the recesses, precision and care, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw