Agent Sandboxing
Code execution gets limited files, network, CPU, memory and scoped credentials; the sandbox must constrain the capability, not just the process tree.
Frame the problem
Security starts with a concrete asset, attacker capability and trust crossing.
Why the system fails
The sandbox can read the workspace, call arbitrary networks or inherit a broad token, so process isolation does not constrain real authority.
The important question is not “what is Agent Sandboxing?” but “which assumption let untrusted data or an over-scoped identity cross agent-generated workload → host and infrastructure?” Trace the decision at the boundary, then constrain what can happen after the first control fails.
Design the control in layers
Start with the control closest to the interpretation or privilege boundary: Use ephemeral isolation with default-deny files and network Then add a control that reduces blast radius and telemetry that proves the decision was enforced.
The resulting design is not labelled secure. Record the identified controls, the known failure paths, the remaining exposure, and the evidence you would need during an incident.
| Prevent | Detect | Recover |
|---|---|---|
| Use ephemeral isolation with default-deny files and network · Mount only explicit inputs and outputs · Provide task-scoped short-lived credentials and hard resource limits | Sandbox policy violations and unexpected egress/resource use | Contain the affected identity or component, scope impact from audit evidence, and preserve a regression test. |
Key points
- Asset: Host, repositories, secrets, internal services and compute budget.
- Boundary: Agent-generated workload → host and infrastructure
- Primary control: Use ephemeral isolation with default-deny files and network
- Detection signal: Sandbox policy violations and unexpected egress/resource use
- Always ask what limits damage when the primary control fails.
Boundary control exercise
This lesson uses the shared boundary-control exercise.
Follow the attack
Safe conceptual simulation: capability → missing control → crossed boundary → asset impact.
- 1Attacker starts with: Malicious instructions or generated code executed by the agent.
- 2The sandbox can read the workspace, call arbitrary networks or inherit a broad token, so process isolation does not constrain real authority.
- 3The weak or missing boundary control is crossed: Agent-generated workload → host and infrastructure
- 4Impact: Secret leakage, destructive changes, lateral movement or cost abuse.
- Secret leakage, destructive changes, lateral movement or cost abuse.
Defend, detect, recover
One prevention is a single point of security failure. Layer it and make failure observable.
- • Use ephemeral isolation with default-deny files and network
- • Mount only explicit inputs and outputs
- • Provide task-scoped short-lived credentials and hard resource limits
- • Sandbox policy violations and unexpected egress/resource use
- • Contain the affected identity or component.
- • Scope access from audit evidence.
- • Fix the boundary and add a regression test.
- • Misconfiguration and new access paths can bypass the intended control.
- • A privileged insider or compromised control plane may still reach the asset.