← All packs

Regression testing for AI agent behavior — why passing CI doesn't mean the agent didn't get worse

You tweak a system prompt, swap a model version, or change which tools an agent can reach. Unit tests still pass, because unit tests check whether code paths execute correctly, not whether the agent's judgment got better or worse. Two weeks later someone notices the agent now asks three clarifying questions when it used to ask one, or stopped using a tool it should have used, or started writing longer commit messages nobody wanted. Nothing broke in the sense CI checks for. Something broke in the sense that actually matters. Closing that gap needs a different kind of test: a regression eval over behavior, run the same way a test suite is run — before every change ships, not after a user complains.

A unit test checks output; an eval checks judgment

A unit test asserts that parse(input) == expected, and that assertion is either true or false, every time. An agent's response to a prompt is not that kind of thing — the same input can produce differently worded outputs that are equally good, and outputs with matching wording that make different tool calls. Testing agent behavior means grading the parts that matter (did it call the right tool, did it stop when it should have, did it flag the risk it was supposed to flag) while ignoring the parts that don't (exact phrasing, ordering of equivalent steps). Treat every eval as a rubric applied to a transcript, not a string comparison against one.

Build the set from real failures and real successes, not hypotheticals

A golden task set invented in the abstract tends to test what's easy to write a task for, not what actually matters in production. The better source is your own history: every time the agent did something impressively right, or embarrassingly wrong, save the input and the transcript. That becomes a fixed regression case — the "impressively right" ones must keep passing (a fix elsewhere shouldn't quietly undo a past win), and the "embarrassingly wrong" ones become the exact regressions you're checking never come back. A twenty-case set built this way, refreshed whenever something notable happens, catches more real regressions than a two-hundred-case set of made-up scenarios.

Weak eval practiceStronger equivalent
grading by exact string match against a saved "golden" responsegrading by rubric — which facts, tool calls, or constraints must appear, regardless of wording
running the eval suite manually before a big releaserunning it on every prompt, tool, or model change, the same trigger as a test suite
a single pass/fail number for the whole suiteper-case results plus a diff against the last baseline, so one regressed case doesn't hide in an aggregate
writing eval cases only for happy-path tasksincluding known failure modes and edge cases pulled from real transcripts, not just successes
one grading pass with no variance checkrunning non-deterministic cases more than once and grading the distribution, since a single lucky or unlucky sample isn't a reliable signal

Grading is the hard part — pick a method that fits the case

Some checks can stay purely mechanical: did the agent call the tool it should have, did it avoid the one it shouldn't, did a required string or file appear in the output. Those are cheap and deterministic — grade them with code, not judgment. Cases about tone, completeness, or reasoning quality resist code-based grading and often end up graded by another model against a rubric, or by a human on a sample. Whichever you use, keep the grading method fixed across runs — swapping how a case is graded at the same time as the thing being tested makes a regression indistinguishable from a change in the grader.

A regression is a comparison, not a score

An eval score in isolation ("83% pass rate") tells you little; what matters is whether it moved from the last known-good baseline, and on which specific cases. Store the baseline transcripts and scores alongside the eval suite itself, versioned the same way code is, and report new runs as a diff: cases that flipped from pass to fail, cases that flipped fail to pass, and cases whose graded quality moved without flipping a binary outcome. A change that improves the average while quietly breaking one previously-solid case is exactly the failure mode this is meant to catch, and an aggregate score alone will not surface it.

The test to apply before shipping any prompt, tool, or model change: if this made the agent measurably worse at three of the twenty things it used to be good at, would today's checks tell you before a user does? If the honest answer is "only if someone happens to try those three things," the eval suite has a gap, not the change.

Related: context window management for long-running agents — a golden task set only helps if the constraints it's checking are still legible to the agent, not buried under stale tool output by the time it matters.

Related: passing this eval suite is what earns a change the right to reach production at all — it doesn't decide how much production traffic sees it first. See canary and shadow rollouts for agent changes for the rollout stage that catches whatever the eval set didn't think to cover.