Sandboxing and Egress Control
Contain Execution Before You Trust Behaviour
Any agent that runs code, executes shell commands, or processes untrusted files needs an execution container with explicitly declared boundaries, and the declaration should be positive rather than negative — state what it may reach, not what it may not. Filesystem access limited to a working directory populated with exactly the inputs the task needs. No credentials mounted into the environment, because a secret present in a sandbox is a secret available to anything the sandbox runs. Resource limits on CPU, memory, processes and wall-clock, so a runaway or deliberately expensive workload terminates rather than consuming the host. Fresh state per run, so nothing persists between tasks unless a deliberate mechanism carries it. Containment is what makes the rest of the design tolerable: it converts arbitrary code execution from an incident into a contained event with a known boundary.
- Declare what the sandbox may reach, not what it may not
- No credentials inside the sandbox — presence equals availability
- Resource and time limits handle both accidents and deliberate exhaustion
- Fresh state per run unless persistence is an explicit, reviewed feature
Egress Is the Control That Holds
Default-deny egress deserves its reputation as the highest-leverage control in this space, for a specific reason: it is enforced by infrastructure rather than by the model or the application, so it holds even when everything above it has been fully persuaded. Deny all outbound connections and allowlist the specific destinations the task requires. Prefer an authenticated forward proxy that can log and constrain requests over network-level rules alone, because the proxy sees intent — full URLs, methods, sizes — where a firewall sees addresses. Control name resolution too, since resolver traffic is an outbound channel in its own right. And treat the allowlist as a security artefact under change control: allowlists widen quietly, one urgent exception at a time, and an allowlist that includes a general-purpose service that can relay arbitrary content is functionally open.
- Enforced in infrastructure, so it survives a fully persuaded model
- Deny by default; allowlist specific destinations for specific tasks
- Prefer a logging proxy over address-level rules — it sees intent, not just endpoints
- Constrain name resolution; an allowlist entry that can relay arbitrary content is an open door
Where Sandboxing Does Not Help
Be precise about what containment does not address, because over-claiming here leads to a false sense of completion. A sandbox constrains the code the agent runs; it does nothing about the tools the agent legitimately calls outside it. If the agent has an approved tool that sends messages, a perfect sandbox is irrelevant to a message being sent with attacker-chosen content. It does not stop the model from producing wrong or harmful output, since the output leaves the sandbox by design. It does not stop data reaching the model's context and then leaving through a permitted destination. And a sandbox with mounted credentials or a broad allowlist is a sandbox in name only. Sandboxing addresses one specific failure — arbitrary execution reaching the host or the network — and it should be described that way in the threat model rather than as generalised containment.
- Sandboxes contain executed code, not the agent's legitimate tool calls
- Harmful output leaves the sandbox by design — that is what output is
- Mounted credentials or a wide allowlist defeat the whole control
- Record precisely which threat the sandbox addresses so nobody over-credits it
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.