The AI Learning Hub Journal

Sandboxing and Egress Control

State what it may reach — deny everything elsecontainment converts arbitrary execution into a contained event; egress control holds when everything above it is persuadedSANDBOX — DECLARED POSITIVELYexecuted code,untrusted filesno credentials insidepresence = availabilityworking directory holdsexactly the task's inputsCPU · memory · process ·wall-clock limitsfresh state per run —nothing persistsdeclared positively — what itmay reach, not what it may notauthenticated forward proxysees intent: full URLs,methods, sizes — and logscontrols name resolutionallowlisteddestinationseverything elsedefault denythe allowlist is a security artefact underchange control — one relay-capable entrymakes it functionally openWHAT THE SANDBOX DOES NOT ADDRESS — SAY SO IN THE THREAT MODELlegitimate tool callsan approved send-messagetool sits outside itharmful outputoutput leaves the sandboxby designpermitted-path leaksdata can exit through anallowlisted destinationself-defeatmounted creds or a wideallowlist void the controlit addresses one failure — arbitrary execution reaching the host or network — credit it for exactly thatEGRESS CONTROL IS ENFORCED BELOW THE MODEL — IT HOLDS WHEN PERSUASION SUCCEEDSdeny outbound by default, allowlist per task, and prefer the proxy whose log your detection will read
Declare positively what execution may reach, deny egress by default through a logging proxy, and state plainly what the sandbox does not cover.

Contain Execution Before You Trust Behaviour

Any agent that runs code, executes shell commands, or processes untrusted files needs an execution container with explicitly declared boundaries, and the declaration should be positive rather than negative — state what it may reach, not what it may not. Filesystem access limited to a working directory populated with exactly the inputs the task needs. No credentials mounted into the environment, because a secret present in a sandbox is a secret available to anything the sandbox runs. Resource limits on CPU, memory, processes and wall-clock, so a runaway or deliberately expensive workload terminates rather than consuming the host. Fresh state per run, so nothing persists between tasks unless a deliberate mechanism carries it. Containment is what makes the rest of the design tolerable: it converts arbitrary code execution from an incident into a contained event with a known boundary.

  • Declare what the sandbox may reach, not what it may not
  • No credentials inside the sandbox — presence equals availability
  • Resource and time limits handle both accidents and deliberate exhaustion
  • Fresh state per run unless persistence is an explicit, reviewed feature

Egress Is the Control That Holds

Default-deny egress deserves its reputation as the highest-leverage control in this space, for a specific reason: it is enforced by infrastructure rather than by the model or the application, so it holds even when everything above it has been fully persuaded. Deny all outbound connections and allowlist the specific destinations the task requires. Prefer an authenticated forward proxy that can log and constrain requests over network-level rules alone, because the proxy sees intent — full URLs, methods, sizes — where a firewall sees addresses. Control name resolution too, since resolver traffic is an outbound channel in its own right. And treat the allowlist as a security artefact under change control: allowlists widen quietly, one urgent exception at a time, and an allowlist that includes a general-purpose service that can relay arbitrary content is functionally open.

  • Enforced in infrastructure, so it survives a fully persuaded model
  • Deny by default; allowlist specific destinations for specific tasks
  • Prefer a logging proxy over address-level rules — it sees intent, not just endpoints
  • Constrain name resolution; an allowlist entry that can relay arbitrary content is an open door

Where Sandboxing Does Not Help

Be precise about what containment does not address, because over-claiming here leads to a false sense of completion. A sandbox constrains the code the agent runs; it does nothing about the tools the agent legitimately calls outside it. If the agent has an approved tool that sends messages, a perfect sandbox is irrelevant to a message being sent with attacker-chosen content. It does not stop the model from producing wrong or harmful output, since the output leaves the sandbox by design. It does not stop data reaching the model's context and then leaving through a permitted destination. And a sandbox with mounted credentials or a broad allowlist is a sandbox in name only. Sandboxing addresses one specific failure — arbitrary execution reaching the host or the network — and it should be described that way in the threat model rather than as generalised containment.

  • Sandboxes contain executed code, not the agent's legitimate tool calls
  • Harmful output leaves the sandbox by design — that is what output is
  • Mounted credentials or a wide allowlist defeat the whole control
  • Record precisely which threat the sandbox addresses so nobody over-credits it

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.