Permission scoping bounds which tools an agent can call directly. It says nothing about what happens when one of those tools is "run this script" and the script is something the agent just wrote. At that point the risk isn't the agent's tool call anymore — it's whatever the generated code does once it's executing, and no amount of scoping the invocation itself changes that the code inside can do anything the process it runs in is able to do.
A human who writes a script usually has some model of what it does before running it. An agent that writes a script to solve a subtask and immediately executes it is skipping that step by construction — the first time the code's actual behavior is observed is during execution, not before it. Most of the time this is harmless: the code does what it was meant to do. But "most of the time" is exactly the population a sandbox is for — the run where a path is wrong, a dependency does something unexpected, or the code the agent wrote was itself shaped by injected instructions from content it read earlier in the same task. None of these require the agent to be malicious; they only require the code to be untrusted, which generated-and-immediately-run code always is.
A sandbox that relies on the generated code being well-behaved isn't a sandbox — it's a hope. The boundary that holds regardless of what the code contains is one enforced by the OS or runtime beneath it: a container or VM with its own filesystem view, a restricted user with no write access outside an explicit scratch directory, no ambient credentials in its environment. This is the same logic as keeping secrets out of a shell an agent can read, aimed one layer deeper: code the agent wrote shouldn't be able to read the agent's own credentials either, even though it's running on the agent's behalf.
| Risky pattern | Contained equivalent |
|---|---|
| generated script runs as the same user/process as the agent | runs in a separate container/VM or restricted user with its own filesystem root |
| script has the agent's full working directory writable | script gets a narrow scratch directory; the rest of the tree is read-only or absent |
| script has unrestricted outbound network access | no network egress by default; specific hosts allow-listed only when the task requires them |
| script inherits the agent's environment variables, including credentials | script gets an explicit, minimal environment with no ambient secrets |
| no limit on runtime, memory, or CPU | hard timeout and resource ceiling, so a runaway or hung process fails bounded instead of consuming the host |
Filesystem isolation gets attention because "the script deleted files" is a vivid failure. Network isolation gets skipped more often, and it's the one that turns a contained mistake into an exfiltration path: code that can read a local file and also make an outbound request can combine the two, deliberately or by following an injected instruction. Closing default network egress and allow-listing specific hosts only when a task genuinely needs them removes that combination regardless of why the code tried it.
A hard wall-clock timeout, a memory ceiling, and a process/thread limit on generated-code execution aren't there to make things fast — they're there so that a script with an infinite loop, a fork bomb, or a memory leak fails as a killed process instead of degrading the host the agent itself is running on. Without a ceiling, the failure mode of "the generated code had a bug" and the failure mode of "the host became unresponsive" are the same event.
The question to ask before an agent's next "let me write a quick script for this": if that script did something completely different from what it claims — read a file it has no reason to touch, or opened a connection to a host with no relation to the task — what would actually stop it? If the honest answer is "nothing, it's running in the same place as everything else," the fix is a process boundary, not a more carefully worded prompt.