← All packs

Designing feedback memory for agents — capturing corrections and confirmations, not just facts

An agent that gets corrected — "don't mock the database in these tests," "stop summarizing at the end of every response" — and forgets it by the next session will get corrected again, for the same thing, indefinitely. That failure is obvious enough that most agent memory designs at least attempt to fix it: save the correction, recall it next time. What's less obvious, and gets skipped almost as often as it gets built, is the other half: when a user accepts an unusual approach without objecting, or explicitly says "yes, exactly, keep doing that," that's feedback too — just quieter. A feedback memory system tuned to catch only corrections produces an agent that avoids its old mistakes but slowly drifts away from judgment calls the user already validated, because nothing ever recorded that they were validated.

Why this isn't the same problem as persistent memory

Persistent memory, covered separately, stores facts about the world: the user's role, a business constraint, a decision that was made. Feedback memory stores something narrower and more behavioral: guidance about how the agent should work, extracted from the user's reactions to what it already did. The distinction matters because the two decay differently and get invalidated by different things. A fact goes stale when the world changes — a function gets renamed, a person changes teams. A piece of feedback goes stale when its scope was narrower than it looked — a "don't do X" said about one specific file gets over-applied to every file, or a confirmed approach for one kind of refactor gets assumed to hold for a kind it was never tested against. Storing both in one undifferentiated pile makes it hard to apply either kind of staleness check correctly.

Corrections are easy to catch; confirmations are not

A correction announces itself: the user says "no, not that," "stop doing X," or otherwise directly interrupts an approach in progress. It has a clear trigger and a clear moment to capture it. A confirmation has no equivalent flag most of the time — it's the absence of a correction where one might have been expected, or a short affirmative ("yes exactly," "perfect, keep doing that") that's easy to read as small talk and move past. The asymmetry means a feedback system built by instinct, without a deliberate rule for the quiet case, ends up correction-only by default — not because confirmations were judged unimportant, but because nobody had to design a mechanism to notice them, while the correction handling built itself.

SignalEasy to miss becauseWhat to capture
Direct correction ("don't do X")It doesn't get missed — it interrupts the flowThe rule, plus why it was given if stated (a past incident, a stated preference)
Explicit confirmation ("yes, exactly")Reads as filler acknowledgment rather than a data pointThe approach just taken, framed as validated — not just "not wrong" but "keep doing this"
Silent acceptance of an unusual choiceNo signal at all unless you're watching for the absence of pushback on something non-defaultWeaker than an explicit confirmation — worth a tentative note, not a firm rule, until it recurs
Repeated correction of the same thingEach instance looks like an isolated one-off if not compared against prior entriesAn update to the existing entry with wider scope, not a duplicate — signals the first version was scoped too narrowly

The "why" is what makes an old entry usable

"Don't mock the database in tests" is an instruction. "Don't mock the database in tests, because a mocked pass once hid a broken migration that only failed against the real schema" is a judgment an agent can extend to a case the original correction never mentioned — a new test file, a different table. Without the reason, a feedback entry is a rule to pattern-match against literally, which breaks the moment a new situation resembles the corrected one only partially. With the reason, edge cases become answerable: does this new case share the underlying risk the correction was guarding against, or does it just superficially resemble the old one. The same applies to confirmations — "the single bundled PR was right because splitting would've been pure churn" tells a future agent when to reuse that judgment and when the situation is different enough not to.

Scope creep in both directions

A feedback entry written too narrowly gets treated as a one-off and the same correction happens again in a slightly different context. Written too broadly, it gets over-applied to situations the user never actually weighed in on — an agent that was told not to add speculative error handling in one function starts refusing all error handling everywhere, including at genuine system boundaries where it belongs. Neither failure is solved by writing more text; it's solved by explicitly recording the boundary the feedback was given inside — "this applies to internal call sites, not to user-facing API boundaries" — at the time the feedback is saved, while the context that produced it is still available, rather than trying to reconstruct the intended scope later from a bare rule.

The test to apply when saving feedback: could a future agent, reading only this entry with none of today's conversation, tell not just what to do but why, and where the boundary of "this applies" sits? If the answer only covers the what, the entry will eventually misfire in one direction or the other.

Related: persistent memory design covers the adjacent problem of storing facts rather than behavioral guidance, and shares the same staleness discipline — recorded then, not necessarily true now.

See also prompt versioning for the case where feedback is stable and general enough to graduate out of a memory entry entirely and get baked directly into the system prompt or CLAUDE.md — at which point it stops needing to be recalled because it's simply always there.