← All packs

Parallel tool calls for agents — which ones are safe to fire together

Most agent harnesses let a model request several tool calls in a single turn and run them concurrently — read three files at once, hit two independent APIs at once, check the status of several unrelated jobs at once. When the calls really are independent, this is close to free speed: the wall-clock cost drops from the sum of the calls to roughly the slowest one. The trap is that "no obvious dependency between these two calls" is not the same claim as "these two calls are safe to run at the same time," and treating the first as proof of the second is where parallel tool-calling bugs come from. Two calls with completely disjoint arguments can still collide on something neither call's signature mentions: a file they both touch, a rate limit they both draw from, or an ordering assumption one of them silently depends on.

Why this isn't caching or rate-limit backoff

Caching decides whether to make a call at all, by checking if a previous result is still valid (see tool call caching); parallelization assumes every call is going to happen and asks only whether several of them can happen at the same moment without interfering. Rate-limit backoff (see rate limits and backoff) is about what an agent does after a call fails or gets throttled; parallelization is upstream of that — it's what turns three calls that would each have succeeded alone into a burst that trips the same limit or a race that corrupts shared state. A tool call can be perfectly idempotent and still be unsafe to run in parallel with another call, because idempotence is about repeating one call, not about two different calls landing at the same instant.

Disjoint arguments don't mean disjoint effects

The instinct is to check whether two tool calls share an argument — two file reads with different paths look obviously independent. But "independent" has to be judged against everything the calls actually touch, not just their argument lists. A call that appends a line to a shared log file and a call that reads that same log file's line count take different arguments (a string to append versus nothing) but are not safe to run concurrently, because one can observe the other mid-write. A call that increments a counter in a database and a second call doing an unrelated update to a different row in the same table might look independent by primary key, but if both go through a connection pool with a hard cap, running them at once doesn't corrupt data — it just means one of them silently queues behind the other while the agent's mental model assumes both ran at full speed.

Looks independent becauseActually sharesWhat goes wrong in parallel
Different file pathsSame directory being listed by a third callThe listing can be taken mid-write, missing or double-counting a file being created
Different API endpoints, same providerA shared per-account rate limitBoth calls succeed individually but the pair trips a limit neither would alone
Different order IDsThe same inventory count they both decrementClassic lost-update race: both read the same stock number before either writes it back
A read and a write with no shared argumentThe read depends on the write having already happened (an implicit ordering, not stated anywhere)The read runs first by chance and returns pre-write data, and nothing in either call's signature flagged that dependency

Ordering dependencies are often implicit, not declared

The riskiest case isn't two calls that touch the same resource — that's at least visible if someone looks for it. It's two calls where one is supposed to happen after the other for reasons that live outside the tool's interface: create-a-resource-then-reference-it, write-a-config-then-reload, provision-a-thing-then-check-its-status. Nothing about the tool signatures says "call B depends on call A having completed," so a model that reasons purely from "do these look independent" will happily fire them together and get whichever ordering the scheduler happens to produce. This has to be handled either by the tool layer refusing to accept a reference to something that doesn't exist yet, or by the agent (or a plan-then-execute step ahead of it) explicitly grouping dependent calls into sequential stages rather than one parallel batch.

Let the tool declare its own safety class instead of guessing from outside

The most reliable fix isn't a smarter model, it's a tool interface that states its own concurrency contract instead of leaving the agent to infer it. A tool definition that includes a field like "concurrency_safe": true for pure reads with no shared mutable state, and false — or a "conflicts_with": ["inventory.update"] list — for anything that writes shared state, lets the harness enforce the boundary mechanically: batch the safe ones, serialize anything flagged unsafe or in conflict, regardless of what the model's tool-call plan for that turn looked like. This moves the correctness question from "did the agent reason about this correctly, every time, for every tool combination" to "did the tool author declare the contract correctly, once" — a much smaller and more auditable surface.

The test to apply: for any two tool calls an agent is about to fire together, ask not "do their arguments overlap" but "is there anything in the world — a file, a row, a counter, a limit, an assumption about what already ran — that both calls touch." If the honest answer is "I'm not sure," that pair belongs in sequence until someone checks, not in a parallel batch on the strength of a hunch.

A parallel batch that corrupts shared state is often *also* not retriable the normal way — see idempotent tool design for why a retry after a partial-parallel failure can double an effect instead of fixing it.