← All packs

Sandboxing code an agent writes and runs — isolation, not just permission scope

Permission scoping bounds which tools an agent can call directly. It says nothing about what happens when one of those tools is "run this script" and the script is something the agent just wrote. At that point the risk isn't the agent's tool call anymore — it's whatever the generated code does once it's executing, and no amount of scoping the invocation itself changes that the code inside can do anything the process it runs in is able to do.

Generated code is not reviewed code, even when it looks fine

A human who writes a script usually has some model of what it does before running it. An agent that writes a script to solve a subtask and immediately executes it is skipping that step by construction — the first time the code's actual behavior is observed is during execution, not before it. Most of the time this is harmless: the code does what it was meant to do. But "most of the time" is exactly the population a sandbox is for — the run where a path is wrong, a dependency does something unexpected, or the code the agent wrote was itself shaped by injected instructions from content it read earlier in the same task. None of these require the agent to be malicious; they only require the code to be untrusted, which generated-and-immediately-run code always is.

The isolation boundary that actually matters is the process, not the intent

A sandbox that relies on the generated code being well-behaved isn't a sandbox — it's a hope. The boundary that holds regardless of what the code contains is one enforced by the OS or runtime beneath it: a container or VM with its own filesystem view, a restricted user with no write access outside an explicit scratch directory, no ambient credentials in its environment. This is the same logic as keeping secrets out of a shell an agent can read, aimed one layer deeper: code the agent wrote shouldn't be able to read the agent's own credentials either, even though it's running on the agent's behalf.

Risky patternContained equivalent
generated script runs as the same user/process as the agentruns in a separate container/VM or restricted user with its own filesystem root
script has the agent's full working directory writablescript gets a narrow scratch directory; the rest of the tree is read-only or absent
script has unrestricted outbound network accessno network egress by default; specific hosts allow-listed only when the task requires them
script inherits the agent's environment variables, including credentialsscript gets an explicit, minimal environment with no ambient secrets
no limit on runtime, memory, or CPUhard timeout and resource ceiling, so a runaway or hung process fails bounded instead of consuming the host

Network egress is the default to close, not the default to leave open

Filesystem isolation gets attention because "the script deleted files" is a vivid failure. Network isolation gets skipped more often, and it's the one that turns a contained mistake into an exfiltration path: code that can read a local file and also make an outbound request can combine the two, deliberately or by following an injected instruction. Closing default network egress and allow-listing specific hosts only when a task genuinely needs them removes that combination regardless of why the code tried it.

Resource limits are a safety property, not a performance one

A hard wall-clock timeout, a memory ceiling, and a process/thread limit on generated-code execution aren't there to make things fast — they're there so that a script with an infinite loop, a fork bomb, or a memory leak fails as a killed process instead of degrading the host the agent itself is running on. Without a ceiling, the failure mode of "the generated code had a bug" and the failure mode of "the host became unresponsive" are the same event.

The question to ask before an agent's next "let me write a quick script for this": if that script did something completely different from what it claims — read a file it has no reason to touch, or opened a connection to a host with no relation to the task — what would actually stop it? If the honest answer is "nothing, it's running in the same place as everything else," the fix is a process boundary, not a more carefully worded prompt.