Most agent harnesses let a model request several tool calls in a single turn and run them concurrently — read three files at once, hit two independent APIs at once, check the status of several unrelated jobs at once. When the calls really are independent, this is close to free speed: the wall-clock cost drops from the sum of the calls to roughly the slowest one. The trap is that "no obvious dependency between these two calls" is not the same claim as "these two calls are safe to run at the same time," and treating the first as proof of the second is where parallel tool-calling bugs come from. Two calls with completely disjoint arguments can still collide on something neither call's signature mentions: a file they both touch, a rate limit they both draw from, or an ordering assumption one of them silently depends on.
Caching decides whether to make a call at all, by checking if a previous result is still valid (see tool call caching); parallelization assumes every call is going to happen and asks only whether several of them can happen at the same moment without interfering. Rate-limit backoff (see rate limits and backoff) is about what an agent does after a call fails or gets throttled; parallelization is upstream of that — it's what turns three calls that would each have succeeded alone into a burst that trips the same limit or a race that corrupts shared state. A tool call can be perfectly idempotent and still be unsafe to run in parallel with another call, because idempotence is about repeating one call, not about two different calls landing at the same instant.
The instinct is to check whether two tool calls share an argument — two file reads with different paths look obviously independent. But "independent" has to be judged against everything the calls actually touch, not just their argument lists. A call that appends a line to a shared log file and a call that reads that same log file's line count take different arguments (a string to append versus nothing) but are not safe to run concurrently, because one can observe the other mid-write. A call that increments a counter in a database and a second call doing an unrelated update to a different row in the same table might look independent by primary key, but if both go through a connection pool with a hard cap, running them at once doesn't corrupt data — it just means one of them silently queues behind the other while the agent's mental model assumes both ran at full speed.
| Looks independent because | Actually shares | What goes wrong in parallel |
|---|---|---|
| Different file paths | Same directory being listed by a third call | The listing can be taken mid-write, missing or double-counting a file being created |
| Different API endpoints, same provider | A shared per-account rate limit | Both calls succeed individually but the pair trips a limit neither would alone |
| Different order IDs | The same inventory count they both decrement | Classic lost-update race: both read the same stock number before either writes it back |
| A read and a write with no shared argument | The read depends on the write having already happened (an implicit ordering, not stated anywhere) | The read runs first by chance and returns pre-write data, and nothing in either call's signature flagged that dependency |
The riskiest case isn't two calls that touch the same resource — that's at least visible if someone looks for it. It's two calls where one is supposed to happen after the other for reasons that live outside the tool's interface: create-a-resource-then-reference-it, write-a-config-then-reload, provision-a-thing-then-check-its-status. Nothing about the tool signatures says "call B depends on call A having completed," so a model that reasons purely from "do these look independent" will happily fire them together and get whichever ordering the scheduler happens to produce. This has to be handled either by the tool layer refusing to accept a reference to something that doesn't exist yet, or by the agent (or a plan-then-execute step ahead of it) explicitly grouping dependent calls into sequential stages rather than one parallel batch.
The most reliable fix isn't a smarter model, it's a tool interface that states its own concurrency contract instead of leaving the agent to infer it. A tool definition that includes a field like "concurrency_safe": true for pure reads with no shared mutable state, and false — or a "conflicts_with": ["inventory.update"] list — for anything that writes shared state, lets the harness enforce the boundary mechanically: batch the safe ones, serialize anything flagged unsafe or in conflict, regardless of what the model's tool-call plan for that turn looked like. This moves the correctness question from "did the agent reason about this correctly, every time, for every tool combination" to "did the tool author declare the contract correctly, once" — a much smaller and more auditable surface.
The test to apply: for any two tool calls an agent is about to fire together, ask not "do their arguments overlap" but "is there anything in the world — a file, a row, a counter, a limit, an assumption about what already ran — that both calls touch." If the honest answer is "I'm not sure," that pair belongs in sequence until someone checks, not in a parallel batch on the strength of a hunch.
A parallel batch that corrupts shared state is often *also* not retriable the normal way — see idempotent tool design for why a retry after a partial-parallel failure can double an effect instead of fixing it.