← All packs
Agent Engineering Reference
45 free, practical guides on building reliable AI coding agents — tool design, permissions, memory, evals, observability, cost control, and failure modes. Written for practitioners; no signup, no paywall.
Machine-readable index: guides-index.json (CC-BY-4.0).
Guides
- What to log when an agent acts autonomously — building a reconstructable audit trail — A human's mistake is explainable because you can ask them. An agent's mistake is explainable only if you logged enough to reconstruct the de
- Canary and shadow rollouts for agent changes — shipping a new prompt or model without betting all your traffic — A prompt or model change that passed every offline eval can still regress in production, because production traffic is never exactly the eva
- Sandboxing code an agent writes and runs — isolation, not just permission scope — Permission scoping limits which tools an agent can call. It doesn't limit what code the agent writes and then executes can do once it's runn
- Context window management for long-running agents — what to keep, what to drop, what to summarize — A long agent session doesn't fail because the model got dumber — it fails because the context window filled up with stale tool output and th
- Cost and token budgets for autonomous agents — stopping a runaway loop before the bill does — An autonomous agent has no innate sense of cost. A practical approach to per-task token budgets, detecting runaway loops before they burn th
- Regression testing for AI agent behavior — why passing CI doesn't mean the agent didn't get worse — A prompt tweak, model swap, or tool change can silently make an agent worse at the thing it was actually good at, while every unit test stay
- Designing feedback memory for agents — capturing corrections and confirmations, not just facts — Persistent memory stores facts about the world. Feedback memory stores something different: guidance the user gave about how to work, both c
- Designing when an agent should stop and ask a human — escalation without alert fatigue — Permission scoping decides what an agent is allowed to do. Escalation design decides when it should stop and ask a human anyway. A practical
- Scoping an agent's permissions — why "can run any command" is the bug — A CI permission guide covers who's watching. A secrets guide covers what leaks. This one covers a third axis: how wide the agent's tool gran
- Designing persistent memory for agents — what survives across sessions, what doesn't — Context window management decides what stays inside one session. Persistent memory decides what an agent should still know a week later, in
- Prompt injection defense for agents that read untrusted content — Permission scoping limits what an agent can do; secrets handling limits what it can leak. Neither stops a webpage, issue comment, or file th
- Versioning and rolling back agent system prompts — treating instructions like a deployable — A system prompt that changed last Tuesday and broke a workflow is a production incident with no stack trace. How to version, diff, and roll
- Rate limits and backoff for agent tool calls — surviving 429s without cascading failure — An agent that retries a failing API in a tight loop turns one degraded dependency into an outage. How to design backoff, jitter, and circuit
- Secrets and credentials in an AI agent's shell — where they actually leak — An agent with shell access can read, echo, log, and paste any secret in its environment. A checklist for where credentials actually leak in
- Validating and repairing agent structured output — handling malformed JSON before it reaches downstream code — An agent's structured output looks like JSON almost all the time, which is exactly what makes the remaining failures dangerous. A practical
- Supply-chain trust for AI agents — the tools and skills you didn't write — Permission scoping, sandboxing, prompt injection defense, and multi-agent trust boundaries all assume the agent's own tools and skills are t
- Checkpointing long-running agent tasks — resuming without redoing or duplicating work — An agent doing a 40-step migration gets killed at step 23 by a deploy, a timeout, or an operator interrupt. A guide to checkpointing agent t
- Caching tool calls for agents — when memoizing a result is safe and when it lies — An agent that re-reads the same file, re-runs the same search, or re-calls the same API ten times in one session is wasting tokens and time
- Parallel tool calls for agents — which ones are safe to fire together — An agent that can call several tools in the same turn will often fire them together to save time. Some of those calls are genuinely independ
- Tool contract drift — what happens when the API underneath an agent's tool changes shape — A tool an agent has called successfully for months can start failing silently the day its underlying API adds a required field, renames a pa
- Deprecating an agent tool without breaking sessions that are already mid-task — Removing or renaming a tool an agent depends on is safe for new sessions and dangerous for existing ones — a session that planned around a t
- Writing tool descriptions agents actually use correctly — schema and wording that prevent wrong-tool and bad-argument calls — An agent doesn't fail to use a tool because it's a bad tool. It fails because the description didn't say the one thing that would have preve
- Designing tool error messages agents can actually recover from — the runtime half of tool design — A well-described tool can still fail at runtime, and what it returns when it fails decides whether the agent corrects itself in one step or
- Sizing tool output for agents — pagination, truncation, and summarization at the source — A tool that can return a thousand rows or a ten-thousand-line file has to decide what it hands back before the agent ever sees it. How to de
- Tracing an agent's decision path — observability beyond print statements — A stuck or slow agent run produces a wall of chat text, not a stack trace. How to instrument agent sessions with spans and correlation IDs s
- Blameless postmortems: a checklist for the three ways they quietly fail — Most postmortems aren't ruined by naming names. They're ruined by stopping at the trigger instead of the root cause, reconstructing the time
- How to generate a changelog from git log without an LLM — A practical guide to turning Conventional Commits into a categorized changelog and a correct semver bump using plain git log parsing — no AP
- Concurrency bug finding: name the interleaving, not "might race" — Concurrency bugs are code that's individually correct and only breaks under a specific interleaving. A practical checklist for shared mutabl
- Dependency upgrade risk: why the changelog and the semver bump both lie — A dependency's changelog tells you what the maintainer thinks changed. It doesn't tell you what your code actually calls, whether that call
- Diagnosing flaky tests: a checklist for the four real causes — Flaky tests almost always come from one of four mechanisms: timing, shared state, ordering, or an external dependency. A practical checklist
- Error log triage: volume is not impact — A raw error dump is a triage problem, not a read-every-line problem. A practical checklist for grouping errors by real root signature and ra
- Running Claude Code headless in CI — what changes when no one can click "allow" — Claude Code's permission prompts assume a human at a terminal. In CI there isn't one, so a naive headless run just hangs until the job times
- Hook or skill? When automation needs to be guaranteed, not just likely — Hooks and skills both let Claude Code do things without being asked, but only one of them is guaranteed to run. A practical guide to picking
- Designing tools for AI agents to call — idempotency, dry-run, and blast radius — An agent will call your tool more than once for the same intent — on retry, on ambiguity, on a re-run after a crash. A guide to designing ag
- Memory leak hunting: reference chains, not "check for leaks" — Restarting the process fixing it temporarily is the strongest leak signal there is. A practical checklist of the six mechanisms that actuall
- Trust boundaries between agents — when a subagent's output becomes another agent's input — Permission scoping, sandboxing, and prompt injection defense all bound a single agent's contact with untrusted content. Multi-agent systems
- Perf regression hunting: it's a diff problem, not a profiling problem — When something got slower and there's a known-good 'before,' profiling from scratch wastes the one advantage you have: a baseline to diff ag
- Designing reversible actions for autonomous agents — undo, not just retry-safety — Idempotency makes a retry safe. It says nothing about whether the first call was a mistake you can take back. A guide to designing agent act
- Reviewing an AI agent's diff — what actually needs checking before you merge it — An agent's diff passes its own tests and still hides the wrong kind of bug. A checklist for what to actually look at before merging AI-gener
- Root-cause tracing: why the first plausible cause is usually wrong — When a bug is traced backward, the trail almost always stops at the first plausible-looking cause instead of the real origin. A practical ch
- Safe database migrations: a checklist for lock, rollback, and rollout risk — Most migration incidents come from three gaps: a lock the app can't tolerate, a rollback that isn't actually symmetric, and a rolling deploy
- Skill, subagent, or slash command? Picking the right container for a workflow — Claude Code gives you three ways to package a repeatable workflow: a skill, a subagent, or a slash command. A practical decision guide based
- Git worktrees for parallel AI coding agents — one repo, isolated working trees — Running more than one Claude Code agent against the same repo at once causes edit collisions and stale state. Git worktrees give each agent
- Writing a CLAUDE.md that actually changes behavior — Most CLAUDE.md files are ignored past the first few lines. A practical checklist for what to put in project instructions, what to leave to s
- Writing task descriptions an agent can actually execute — tickets, issues, and prompts — A CLAUDE.md sets repo-wide context once. A task description sets the goal for one run. Most agent failures blamed on 'the model didn't under
Also: gitrecap — a free offline changelog generator (no API key, no LLM).