← All packs

Error log triage: volume is not impact

A raw error dump — thousands of lines, dozens of distinct-looking stack traces — invites the wrong first question: "which error appears most?" That's a count, not a priority. The question that actually matters is which root cause, once duplicates are collapsed, is costing the most real impact: users blocked, requests failed, revenue lost. Four checks separate real signal from noise before anyone starts fixing things.

1. Symptom fan-out: one root cause, many stack traces

A single upstream failure — a dependency down, a bad config pushed — surfaces at dozens of different call sites because of retries, wrapper functions, or multiple entry points that all eventually call the same failing thing. Raw line count makes it look like dozens of unrelated problems.

Check: strip outer frames and file/line noise, then group by innermost distinctive frame plus exception type. Count distinct groups, not raw lines. If most of the volume collapses into one group once you do this, you have one root cause, not dozens.

2. Signature collision: same message, different root causes

A generic message — "Connection reset", "Timeout", "NullReferenceException" — gets emitted by many unrelated code paths for unrelated reasons. Grouping by message text alone merges causes that have nothing to do with each other, hiding the fact that one of them is actually the important one.

Check: fingerprint each group by message plus the actual call site or module, not message text alone. Two errors with an identical message but a different origin are two different problems and should never share a bucket.

3. Noise dominating signal: expected conditions logged as errors

Conditions that resolve on their own — a retryable network blip, a validation rejection, a cache miss — get logged at ERROR level instead of INFO or WARN, and their sheer frequency buries the errors that actually need a human.

Check: for each group, ask whether it resolves without action — the retry eventually succeeds, the request completes anyway, no user is left in a bad state. If yes, it's noise regardless of volume, and it should be re-leveled or filtered, not triaged alongside real failures.

4. Volume vs. impact mismatch

One error type can dominate the raw count while affecting a single low-traffic path or one misbehaving client, while a much rarer error is silently blocking the primary checkout or write path for every user who hits it.

Check: rank groups by distinct users or requests actually left in a failed state, not occurrence count. A group with 50 occurrences that all block checkout outranks a group with 5,000 occurrences that all get silently retried and succeed.

Confirm before prescribing

State which group is actually worth fixing first with evidence — the collapsed count, the affected path, the fraction of users blocked — not whichever message appeared most times in the raw dump. If the grouping doesn't clearly separate signal from noise yet, name the specific data that would confirm it (a distinct-user count per group, a check of whether the request ultimately succeeded despite the logged error) rather than shipping a priority order based on a guess. Then say whether the fix addresses the root cause found in check 1, or just silences the symptom at the log line — re-leveling a noisy log is not the same as fixing what's underneath it.

None of this requires an LLM on a small, well-understood system — it's four checks and a grouping pass. Where a model helps is on a large, unfamiliar one: reading thousands of raw log lines, correctly fingerprinting each by real call site rather than surface message text, and cross-referencing which groups actually correlate with failed requests instead of guessing from message frequency, which is what the error-log-triager skill in the Claude Code Debug Pack does. If triage points to a specific bug, see the companion guide on root-cause tracing; if it points to a slowdown instead of a hard failure, see the one on perf regression hunting; if the errors correlate with memory climbing over the process lifetime, see the one on memory leak hunting; if the errors only show up under concurrent load or in a specific ordering, see the one on concurrency bug finding.