You tweak a system prompt, swap a model version, or change which tools an agent can reach. Unit tests still pass, because unit tests check whether code paths execute correctly, not whether the agent's judgment got better or worse. Two weeks later someone notices the agent now asks three clarifying questions when it used to ask one, or stopped using a tool it should have used, or started writing longer commit messages nobody wanted. Nothing broke in the sense CI checks for. Something broke in the sense that actually matters. Closing that gap needs a different kind of test: a regression eval over behavior, run the same way a test suite is run — before every change ships, not after a user complains.
A unit test asserts that parse(input) == expected, and that assertion is either true or false, every time. An agent's response to a prompt is not that kind of thing — the same input can produce differently worded outputs that are equally good, and outputs with matching wording that make different tool calls. Testing agent behavior means grading the parts that matter (did it call the right tool, did it stop when it should have, did it flag the risk it was supposed to flag) while ignoring the parts that don't (exact phrasing, ordering of equivalent steps). Treat every eval as a rubric applied to a transcript, not a string comparison against one.
A golden task set invented in the abstract tends to test what's easy to write a task for, not what actually matters in production. The better source is your own history: every time the agent did something impressively right, or embarrassingly wrong, save the input and the transcript. That becomes a fixed regression case — the "impressively right" ones must keep passing (a fix elsewhere shouldn't quietly undo a past win), and the "embarrassingly wrong" ones become the exact regressions you're checking never come back. A twenty-case set built this way, refreshed whenever something notable happens, catches more real regressions than a two-hundred-case set of made-up scenarios.
| Weak eval practice | Stronger equivalent |
|---|---|
| grading by exact string match against a saved "golden" response | grading by rubric — which facts, tool calls, or constraints must appear, regardless of wording |
| running the eval suite manually before a big release | running it on every prompt, tool, or model change, the same trigger as a test suite |
| a single pass/fail number for the whole suite | per-case results plus a diff against the last baseline, so one regressed case doesn't hide in an aggregate |
| writing eval cases only for happy-path tasks | including known failure modes and edge cases pulled from real transcripts, not just successes |
| one grading pass with no variance check | running non-deterministic cases more than once and grading the distribution, since a single lucky or unlucky sample isn't a reliable signal |
Some checks can stay purely mechanical: did the agent call the tool it should have, did it avoid the one it shouldn't, did a required string or file appear in the output. Those are cheap and deterministic — grade them with code, not judgment. Cases about tone, completeness, or reasoning quality resist code-based grading and often end up graded by another model against a rubric, or by a human on a sample. Whichever you use, keep the grading method fixed across runs — swapping how a case is graded at the same time as the thing being tested makes a regression indistinguishable from a change in the grader.
An eval score in isolation ("83% pass rate") tells you little; what matters is whether it moved from the last known-good baseline, and on which specific cases. Store the baseline transcripts and scores alongside the eval suite itself, versioned the same way code is, and report new runs as a diff: cases that flipped from pass to fail, cases that flipped fail to pass, and cases whose graded quality moved without flipping a binary outcome. A change that improves the average while quietly breaking one previously-solid case is exactly the failure mode this is meant to catch, and an aggregate score alone will not surface it.
The test to apply before shipping any prompt, tool, or model change: if this made the agent measurably worse at three of the twenty things it used to be good at, would today's checks tell you before a user does? If the honest answer is "only if someone happens to try those three things," the eval suite has a gap, not the change.
Related: context window management for long-running agents — a golden task set only helps if the constraints it's checking are still legible to the agent, not buried under stale tool output by the time it matters.
Related: passing this eval suite is what earns a change the right to reach production at all — it doesn't decide how much production traffic sees it first. See canary and shadow rollouts for agent changes for the rollout stage that catches whatever the eval set didn't think to cover.