Skip to content
Apsan Works

How to Evaluate an AI Agent Before You Trust It

You cannot ship what you cannot measure, and “it seemed good when I tried it” is not measurement. A practical method for testing non-deterministic systems.

6 min readUpdated 13 September 2026

Traditional software testing rests on an assumption that does not hold here: the same input produces the same output. Take that away and most of your testing instincts stop working. Assertions get flaky. Snapshots turn to noise. The whole apparatus assumes a determinism you no longer have.

The answer is not to give up on testing. It’s to test differently, measuring distributions instead of asserting equality.

Start with the golden set

Before any agent code, assemble real examples with known-correct outcomes. This is the foundation, and there’s no substitute for it.

The examples should come from historical records of the work being automated, or cases your team has already handled where you know what the right answer was. If the process is entirely new, have a domain expert work through fifty cases by hand. That’s a week of someone’s time, and it’s the highest-leverage week in the project.

Fifty examples is enough to catch obvious regressions. Two hundred gives reasonable confidence. Past around five hundred you’re into diminishing returns unless you’re covering genuinely distinct categories.

What goes in the set matters more than how many. You want the boring middle, but also the ambiguous cases, the malformed inputs, the ones where the correct answer is “I cannot determine this,” and adversarial cases if the system is externally facing.

Choose a grading method per task

Different outputs need different graders. Using the wrong one is a common way to get numbers that don’t mean anything.

Exact match works for classification, routing, and structured extraction, anywhere there’s exactly one right answer. It’s cheap, unambiguous, and the best option whenever it applies.

Structural validation checks the shape rather than the content: does the output parse, does it satisfy the schema, are the required fields present, are the values in range. It’s fast and deterministic, and it catches a surprising share of real failures on its own.

Rubric grading by a model suits open-ended outputs like summaries, drafted messages, or explanations. You define specific criteria and have a model score against them. This correlates reasonably well with human judgement when the rubric is concrete, and degrades badly when it isn’t. A criterion like “is it good” is worse than useless.

Human review stays necessary for a sampled subset, permanently. It’s the only method that catches problems your rubric didn’t anticipate, which is exactly the category you most need to know about.

Outcome measurement is the strongest signal, and the one people skip because it’s slow. Did the downstream thing actually work? Did the extracted data reconcile? Did the message get a reply? Did the human accept the recommendation without editing it? When you can measure this, it beats every proxy you’ve got.

Measure the things that actually break projects

Accuracy is the obvious metric, and it’s rarely the one that kills you.

Failure mode distribution matters more than the aggregate rate. Ten percent wrong means something different depending on whether the errors are scattered randomly or concentrated in one input category. The second is a bug you can fix. The first is a capability limit.

Calibration asks whether the agent’s expressed confidence tracks its actual accuracy. A system that’s right eighty percent of the time and knows which eighty percent is far more useful than one that’s right ninety percent of the time with uniform confidence, because the first one can route its uncertain cases to a human and the second can’t tell you which cases those are.

Cost and latency percentiles matter too. The p50 is comfortable and the p99 is what people complain about, and agent loops have long tails by nature.

Consistency across runs is worth checking directly: run the same input several times. High variance on identical inputs means the system is fragile in ways a single-pass eval won’t show you.

Wire it into the workflow

An eval suite that runs when someone remembers to run it is a suite that doesn’t run.

It should fire automatically in CI on every change to prompts, tools, models, or retrieval, and it should block the merge if accuracy drops below the current baseline. That second part is the discipline that makes the whole thing worth having. Report results as a diff rather than an absolute score. “Improved on 12 cases, regressed on 3, here are the 3” is something a person can act on. “Score: 0.87” is not. And keep it cheap enough to run constantly. A full eval that costs forty dollars and takes an hour will get skipped, so keep a fast subset for every commit and save the full run for overnight.

Shadow mode before switchover

The last step before an agent takes real actions is running it alongside the existing process without letting it act.

Both the agent and the current process handle the same inputs, but only the current process’s output gets used. You compare the two afterward. This gives you a genuine production-distribution measurement instead of an eval-set proxy, and it lets the people who’ll rely on the system watch it work before they have to trust it.

The disagreements are the valuable part. Every case where the agent and the human differ is either a bug or a sign that your definition of correct was less settled than you thought, and both are worth knowing before you cut over. Run it long enough to cover a full cycle of whatever seasonality your work has, then switch over on the numbers rather than on a launch date. The same discipline carries over almost unchanged to replacing an entire legacy system, not just an agent.

What good enough looks like

There’s no universal threshold. It depends on the cost asymmetry.

If a wrong answer is expensive and a delay is cheap, in law, medicine, finance, anything regulated, you want high confidence thresholds and generous escalation to humans, accepting a lower automation rate in exchange for very few wrong actions. If a wrong answer is cheap and throughput is the constraint, internal triage or first-pass categorisation for instance, you can run at much lower confidence, because the correction cost is a person spending ten seconds.

The question is never “is the agent accurate enough.” It’s “is the agent more accurate than the alternative, at a cost that makes sense, with the errors falling somewhere we can absorb them.” Framed that way, the threshold usually becomes obvious.

Evaluation is the first thing we build on every agent engagement. See our approach, or read about why agents stall in production.

Tell us what is slowing you down

A short conversation is usually enough to tell whether this is a build, an automation, or something you should not do at all. We will tell you which.