← All packs

Checkpointing long-running agent tasks — resuming without redoing or duplicating work

A long-running agent task — a multi-file migration, a batch of forty API calls, a multi-stage data backfill — gets interrupted for reasons that have nothing to do with correctness: a deploy restarts the process, a timeout fires, an operator hits stop, the host gets rescheduled. The task wasn't wrong when it stopped; it was just stopped. What happens next depends entirely on whether anything durable recorded how far it got. Without that record, "resume" really means "start over," which is not merely slow — for a task whose steps have side effects, starting over means re-running the side effects too: a second email sent, a second charge issued, a second row inserted. Checkpointing is the design discipline that makes resuming mean "continue from step 24," not "hope step 1 through 23 are safe to repeat."

Why this isn't the same problem as idempotent tool design

Idempotent tool design makes a single retried call safe — call the same tool twice with the same idempotency key and the second call is a no-op or returns the first result. That solves the problem at the level of one call. Checkpointing solves a different problem: a task made of many calls, where the process itself dies mid-sequence and something outside any single call has to know which of the forty already ran. Idempotency keys are one of the tools a checkpoint system uses, but they don't replace it — a task can be built entirely out of idempotent calls and still have no way to know it was on step 23 when it died, and re-drive all forty from the top, relying on idempotency to make the first 23 harmless repeats rather than skipping them outright. That works, but only if every step actually is idempotent and cheap to repeat; checkpointing is what lets you skip them instead of gambling on that.

Why this isn't the same problem as persistent memory or tracing

Persistent memory answers "what should the agent still know next week" — facts, preferences, prior decisions. A checkpoint isn't a fact the agent should recall for judgment; it's task-execution state — an offset into a specific run that should be read once, at resume time, by the same kind of process that's resuming, and otherwise ignored. Tracing answers "where did the time and cost go" and is typically append-only and historical; a checkpoint is mutable, gets overwritten every step, and only the latest value matters. Conflating the three produces predictable bugs: checkpoint data treated as memory bloats context with stale progress markers from finished runs; trace data treated as a checkpoint tries to resume from a historical span that no longer matches current state.

What has to be in the checkpoint

Missing from the checkpointConsequence on resume
Just "step 23 was reached"Resume knows where it stopped but not whether step 23's side effect (the email, the row, the API call) actually completed before the kill — it might re-run a half-finished step and duplicate the effect
Step number but not the step's input/outputResume can't verify it's continuing the same task instance with the same parameters — a re-triggered run with slightly different inputs silently continues from someone else's progress
No record of external side-effect IDs (email message ID, payment ID, created-row ID)Resume can't check "did this already happen" against the external system, so retry-safety depends entirely on that system's own idempotency rather than the task's
Checkpoint written after the side effect instead of atomically with itA kill between "email sent" and "checkpoint saved" makes the resume re-send — the checkpoint has to be at least as durable as, and ordered correctly relative to, the effect it's recording

The minimum durable unit per step is: which step, what its inputs were, what side-effect identifiers it produced (if any), and confirmation the step's write actually landed — not that it started. Writing the checkpoint before a step's side effect happens is what allows the "at-least-once, use the ID to dedupe" pattern; writing it only after is what allows "exactly the completed steps are recorded, nothing more." Both are valid designs — the failure mode is doing neither consistently, so some steps checkpoint before and some after with no way to tell which.

Resume has to re-verify, not just re-read

Reading the checkpoint and jumping to the recorded step assumes the world hasn't changed since the checkpoint was written — often false after a long enough gap. A file the task was mid-edit on may have been touched by something else; a resource the task expected to still exist may have been deleted by an unrelated process; a downstream system the task called may have since changed state on its own. Cheap, targeted re-verification of the state the next step depends on — not a full re-run of prior steps — is what separates a resume that continues correctly from one that silently proceeds on stale assumptions. This is the same instinct as checking a lock is still held before writing, applied to task state instead of a mutex.

Granularity is a real design choice, not a default

Checkpointing after every single tool call is maximally safe and maximally chatty — a durable write per step, which is fine for a forty-step task and expensive for a forty-thousand-step one. Checkpointing only at coarse phase boundaries ("finished ingestion," "finished transformation") is cheap but means a kill mid-phase re-runs the whole phase, which is only acceptable if every step inside that phase is genuinely idempotent or cheap to repeat. The right granularity is set by the cost of re-doing the largest unit of work between checkpoints, not by a fixed rule — a phase whose steps are all idempotent API calls can checkpoint coarsely; a phase that sends irreversible notifications should checkpoint after every one.

The test to apply: kill the task's process at a random point mid-run, on purpose, and resume it. If any side effect happens twice, or any completed step gets silently redone, the checkpoint is missing the state needed to tell "already done" from "not yet done" — usually the external side-effect ID, or the ordering between the effect and the write that records it. If resume can't even tell where it stopped, there's no checkpoint at all, only progress that happens to be inferable from logs.

A checkpoint saved a week ago can also resume into a toolset that's changed since — see deprecating a tool without breaking sessions already mid-task for why a fixed-date removal deadline is the wrong way to retire a tool a paused session might still reference.