I’ve written several times that you can’t tune an agent by vibes, that “done” is an eval bar rather than a spec, and that the quiet danger of a model upgrade is the regression no one notices. All true, and all slightly useless, because none of it tells you what the artifact looks like. An eval set is a concrete thing you build and maintain, and the reason most teams don’t have one isn’t disagreement — it’s that nobody ever showed them the shape of it. So here it is.

What it isn’t
It isn’t a benchmark. Public benchmarks tell you how a model does on someone else’s problem, which is interesting and almost never decision-relevant for you. It also isn’t a unit test suite: unit tests assert exact equality on deterministic code, and most of what an agent produces is a distribution rather than a value.
An eval set is a curated collection of cases with known-good outcomes, chosen because they discriminate — because a system that handles them is meaningfully better than one that doesn’t. Every case earns its place by being able to fail.
Where the cases come from
Not from imagination, and not from random sampling. Random samples are dominated by the easy middle, so they go green and stay green while the things that actually hurt you slip through.
Four sources, in rough order of value. Production incidents — every time the system got something wrong in a way that mattered, that case goes in the set permanently, with the correct outcome recorded. Known-hard inputs — the formats, the edge geometry, the regional variations your domain experts can list from memory if you ask them. Adversarial cases — inputs designed to induce confident wrongness rather than obvious failure. And the boring golden path, enough of it to catch a catastrophic regression, but far less than instinct suggests.
How many
Fewer than you think, and chosen rather than collected. A set of forty well-chosen cases that each discriminate will tell you more than four thousand sampled ones, and — this is the part that decides whether the practice survives — it stays cheap enough to run on every change. An eval set nobody runs because it takes an hour is not an eval set; it’s a document.
Grow it by incident, not by ambition. The set should get bigger every time something breaks, and essentially never otherwise.
The grader is the hard part
Deciding whether an output is right is where most eval efforts quietly collapse. The hierarchy is simple and worth being strict about.
Exact match wherever you can engineer it. If the task produces a structured artifact — a parsed field, a routing decision, a resolved entity — assert on the artifact, not on prose about it. This is the strongest argument for designing agents so their valuable outputs are structured in the first place.
Property assertions next. You often can’t specify the one correct answer but can specify what must be true of any correct one: the total reconciles, the cited span exists in the source, no field was invented. Properties catch a surprising share of real failures.
A model as judge, last and carefully. Sometimes only judgment will do. When you reach for it, the judge must be independent of the system under test — a different harness, a different prompt, ideally a different model — and it must be spot-checked by a human on a rotating sample. A system grading its own homework produces a number that only ever goes up.
What you record
Each case wants an input, an expected outcome or property, a provenance note saying why it exists, and a date. The provenance matters more than it sounds: two years in, somebody will want to delete a case they don’t understand, and “added after the March incident where we paid the wrong party” is the sentence that stops them.
When the eval disagrees with the demo
This is the moment the practice is actually tested, and it happens to everyone. A change makes the demo visibly better and the eval slightly worse. The temptation — overwhelming, because the demo is what people saw — is to decide the eval is being pedantic.
Resist it once and you have an engineering culture. The eval is the memory of every past failure; the demo is a single sample chosen by the person who built the change. If they genuinely conflict, the right response is to work out which case regressed and why, and either fix it or consciously retire the case with a written reason. What you must not do is ship past it quietly, because that’s the move that turns the set back into a document.
The real reason to build one
An eval set isn’t a quality ritual. It’s the only thing that lets you change anything with confidence — swap a model, restructure a prompt, re-route a task to something cheaper — and know within minutes whether you broke something that used to work. Without it, every improvement is a gamble you can’t price, and the rational response to an un-priceable gamble is to stop improving. Teams without evals don’t ship carefully. They stop shipping.
A companion to the six-part series Cheap to Be Wrong — an architecture for agentic systems.
Related: The Model Upgrade That Quietly Breaks You · A Product Playbook for Agentic AI — The Frameworks I Actually Use.




Leave a Reply