A long-running agent task — a multi-file migration, a batch of forty API calls, a multi-stage data backfill — gets interrupted for reasons that have nothing to do with correctness: a deploy restarts the process, a timeout fires, an operator hits stop, the host gets rescheduled. The task wasn't wrong when it stopped; it was just stopped. What happens next depends entirely on whether anything durable recorded how far it got. Without that record, "resume" really means "start over," which is not merely slow — for a task whose steps have side effects, starting over means re-running the side effects too: a second email sent, a second charge issued, a second row inserted. Checkpointing is the design discipline that makes resuming mean "continue from step 24," not "hope step 1 through 23 are safe to repeat."
Idempotent tool design makes a single retried call safe — call the same tool twice with the same idempotency key and the second call is a no-op or returns the first result. That solves the problem at the level of one call. Checkpointing solves a different problem: a task made of many calls, where the process itself dies mid-sequence and something outside any single call has to know which of the forty already ran. Idempotency keys are one of the tools a checkpoint system uses, but they don't replace it — a task can be built entirely out of idempotent calls and still have no way to know it was on step 23 when it died, and re-drive all forty from the top, relying on idempotency to make the first 23 harmless repeats rather than skipping them outright. That works, but only if every step actually is idempotent and cheap to repeat; checkpointing is what lets you skip them instead of gambling on that.
Persistent memory answers "what should the agent still know next week" — facts, preferences, prior decisions. A checkpoint isn't a fact the agent should recall for judgment; it's task-execution state — an offset into a specific run that should be read once, at resume time, by the same kind of process that's resuming, and otherwise ignored. Tracing answers "where did the time and cost go" and is typically append-only and historical; a checkpoint is mutable, gets overwritten every step, and only the latest value matters. Conflating the three produces predictable bugs: checkpoint data treated as memory bloats context with stale progress markers from finished runs; trace data treated as a checkpoint tries to resume from a historical span that no longer matches current state.
| Missing from the checkpoint | Consequence on resume |
|---|---|
| Just "step 23 was reached" | Resume knows where it stopped but not whether step 23's side effect (the email, the row, the API call) actually completed before the kill — it might re-run a half-finished step and duplicate the effect |
| Step number but not the step's input/output | Resume can't verify it's continuing the same task instance with the same parameters — a re-triggered run with slightly different inputs silently continues from someone else's progress |
| No record of external side-effect IDs (email message ID, payment ID, created-row ID) | Resume can't check "did this already happen" against the external system, so retry-safety depends entirely on that system's own idempotency rather than the task's |
| Checkpoint written after the side effect instead of atomically with it | A kill between "email sent" and "checkpoint saved" makes the resume re-send — the checkpoint has to be at least as durable as, and ordered correctly relative to, the effect it's recording |
The minimum durable unit per step is: which step, what its inputs were, what side-effect identifiers it produced (if any), and confirmation the step's write actually landed — not that it started. Writing the checkpoint before a step's side effect happens is what allows the "at-least-once, use the ID to dedupe" pattern; writing it only after is what allows "exactly the completed steps are recorded, nothing more." Both are valid designs — the failure mode is doing neither consistently, so some steps checkpoint before and some after with no way to tell which.
Reading the checkpoint and jumping to the recorded step assumes the world hasn't changed since the checkpoint was written — often false after a long enough gap. A file the task was mid-edit on may have been touched by something else; a resource the task expected to still exist may have been deleted by an unrelated process; a downstream system the task called may have since changed state on its own. Cheap, targeted re-verification of the state the next step depends on — not a full re-run of prior steps — is what separates a resume that continues correctly from one that silently proceeds on stale assumptions. This is the same instinct as checking a lock is still held before writing, applied to task state instead of a mutex.
Checkpointing after every single tool call is maximally safe and maximally chatty — a durable write per step, which is fine for a forty-step task and expensive for a forty-thousand-step one. Checkpointing only at coarse phase boundaries ("finished ingestion," "finished transformation") is cheap but means a kill mid-phase re-runs the whole phase, which is only acceptable if every step inside that phase is genuinely idempotent or cheap to repeat. The right granularity is set by the cost of re-doing the largest unit of work between checkpoints, not by a fixed rule — a phase whose steps are all idempotent API calls can checkpoint coarsely; a phase that sends irreversible notifications should checkpoint after every one.
The test to apply: kill the task's process at a random point mid-run, on purpose, and resume it. If any side effect happens twice, or any completed step gets silently redone, the checkpoint is missing the state needed to tell "already done" from "not yet done" — usually the external side-effect ID, or the ordering between the effect and the write that records it. If resume can't even tell where it stopped, there's no checkpoint at all, only progress that happens to be inferable from logs.
A checkpoint saved a week ago can also resume into a toolset that's changed since — see deprecating a tool without breaking sessions already mid-task for why a fixed-date removal deadline is the wrong way to retire a tool a paused session might still reference.