An agent that writes and runs code can do almost anything a programmer can: parse a spreadsheet, call an API, train a small model, or, if a malicious document tells it to, read environment variables and post them to a stranger's server. Putting that code in a sandbox is the obvious first step, and choosing the isolation technology is a well-covered decision. The harder question is what the sandbox should allow, for this task, for this user, right now, and who decides.
That decision is the policy, and in practice it is where sandboxes fail. Teams build strong isolation and then give every execution the same broad image, the same network access and the same credentials, because deciding per task felt like too much work. This article treats the policy as the product. It shows how to express a policy declaratively, derive it from the task rather than from the model, compile it into enforcement at every boundary the code can touch, handle requests for more access mid-run, gate what comes out, and test that denials actually deny.
It builds on two pages: sandboxing ADK code execution covers the isolation ladder and starting resource limits in detail, and permission boundaries for autonomous agents covers scoping the credentials an agent holds. Here the unit is a single execution of agent-written code.
Isolation is the floor, policy is the boundary
Isolation technology answers one question: if the code inside turns hostile, how hard is it to reach the host? A hardened container shares the host kernel, so a kernel bug is an escape route. gVisor runs containers on a user-space kernel, installed as the runsc runtime, which intercepts system calls before they reach the host. A microVM such as Firecracker, which AWS built to run Lambda functions, gives each sandbox its own guest kernel. Stronger isolation costs startup time and some compatibility, and the right choice depends mostly on whether code from different tenants shares a machine.
None of this decides what the code is allowed to do through the doors you deliberately leave open. The data-exfiltration attack on an agent does not need a kernel exploit. It needs an outbound network connection, a readable credential, or an output channel nobody inspects. Those doors are governed by policy, and a microVM with an open network and a mounted API key is less safe than a plain container with neither.
What a policy contains
A useful policy is a small declarative document, stored and versioned like code, that states every capability one execution receives. Anything not listed is denied. The shape below is illustrative, not a standard, but each field maps to a concrete enforcement point.
policy: analytics-readonly
version: 7
isolation: gvisor # container | gvisor | microvm
image: sandbox/py-analytics@sha256:... # pinned digest, packages pre-installed
limits: {cpus: 1, memory_mb: 1024, pids: 128, wall_seconds: 60, scratch_mb: 256}
inputs: # copied into /in, read-only
- dataset: "${task.dataset_ref}"
network:
default: deny
allow:
- {host: api.fx-rates.internal, methods: [GET], max_response_kb: 512}
credentials: # never enter the sandbox; injected by the proxy
- {host: api.fx-rates.internal, secret: fx-readonly, header: Authorization}
outputs:
allowed_types: [text/csv, image/png, application/json]
max_total_kb: 2048
destination: agent # agent | user | storage-bucket
executions_per_task: 15Three properties make this workable. The image is pinned by digest and contains every package the task type needs, so the code never runs pip install against the internet. Inputs are copied in by reference, so the policy names data, not host paths. And credentials are bound to hosts, not handed to code, which is the subject of a later section.
Deriving policy from the task, never from the model
The most important rule is who writes the policy. The model must never choose its own capabilities, because the model is exactly the component an attacker can influence through a prompt injection in a document, a web page or a tool result. If the model can say "this task needs network access to example.com", so can the attacker.
Instead, the orchestrator derives the policy from facts the model does not control. The task type selects a template. The tenant and user narrow it: which datasets they may read, which internal APIs their role reaches. The task's own inputs narrow it further, so a task about one dataset gets exactly that dataset. The result is hashed and recorded with the task before any code runs.
def compile_policy(task, user, templates, entitlements):
base = templates[task.type] # e.g. analytics-readonly v7
pol = base.copy()
pol.inputs = [d for d in task.dataset_refs if entitlements.can_read(user, d)]
if len(pol.inputs) != len(task.dataset_refs):
raise PolicyError("task references data the user cannot read")
pol.network.allow = [r for r in base.network.allow
if entitlements.can_call(user, r.host)]
pol.limits = base.limits.clamp(user.plan.max_limits)
pol.hash = sha256_canonical(pol)
audit.record(task.id, policy=pol.name, version=pol.version, hash=pol.hash)
return polNotice that compilation can only remove capabilities from the template, never add them. Widening is a separate, explicit path.
Compiling policy into enforcement
A policy only matters if each field is enforced by something the code cannot bypass. For a container-based sandbox, most fields become flags on the container, and the rest become configuration of services outside it.
def docker_args(pol, task_id):
args = [
"docker", "run", "--rm",
"--network", "none", # no interfaces except via the proxy socket below
"--read-only",
"--tmpfs", f"/scratch:rw,size={pol.limits.scratch_mb}m,noexec,nosuid",
"--cap-drop", "ALL",
"--security-opt", "no-new-privileges",
"--pids-limit", str(pol.limits.pids),
"--memory", f"{pol.limits.memory_mb}m",
"--cpus", str(pol.limits.cpus),
"--user", "65534:65534",
"-v", f"{stage_inputs(pol, task_id)}:/in:ro",
"-v", f"{stage_code(task_id)}:/code:ro", # the agent's script, read-only
"-w", "/scratch", # the only writable place
"-v", f"{proxy_socket(task_id)}:/run/egress.sock",
"-e", "HTTPS_PROXY=unix:///run/egress.sock", # illustrative: client shim speaks to the socket
"--label", f"policy_hash={pol.hash}",
]
if pol.isolation == "gvisor":
args += ["--runtime", "runsc"]
return args + [pol.image, "python", "/code/main.py"]The pattern to copy is that the sandbox has no network interface of its own. Its only route out is a per-task channel to an egress proxy that holds that task's allowlist, so a forgotten firewall rule cannot open a path. How you plumb the channel depends on the runtime; a Unix socket with a small client shim is one option, a dedicated network namespace whose only route is the proxy is another. The wall-clock limit is enforced from outside by killing the whole container, not by a timer inside the code.
Enforcement for microVMs has the same shape with different mechanics: the VM gets a virtual network device connected only to the proxy, and inputs arrive as a read-only block device. The policy document does not change, which is the point of keeping it declarative.
The egress proxy and the credential broker
The proxy is where most real attacks are stopped, so it checks more than host names. It resolves names itself, so the sandbox cannot use DNS as a side channel or point a permitted name at a different address. It checks method and path, caps request and response sizes, and counts total bytes per task, because exfiltration through a permitted API is still exfiltration: an attacker can encode data into query strings sent to an allowed host.
def decide(req, pol, usage):
rule = next((r for r in pol.network.allow if r.host == req.host), None)
if rule is None:
return deny("host-not-allowed")
if req.method not in rule.methods:
return deny("method-not-allowed")
if len(req.url) > 2048 or req.body_bytes > rule.get("max_request_kb", 16) * 1024:
return deny("request-too-large")
if usage.bytes_out + req.body_bytes > pol.network.get("max_out_kb", 256) * 1024:
return deny("egress-budget-exhausted")
cred = pol.credential_for(req.host)
if cred:
req.headers[cred.header] = broker.fetch(cred.secret, task=usage.task_id)
return allow(rule)The credential broker is the second half. Code inside the sandbox never sees a secret: it calls the permitted API without authentication, and the proxy adds the header on the way out, using a short-lived credential scoped to that host. A prompt-injected print(os.environ) finds nothing, and a request to a disallowed host carries no secret even if the allowlist were misconfigured. Redirects are the classic gap: the proxy must apply the same decision to every redirect target, or an allowed host becomes a relay.
When the code needs more
Legitimate tasks sometimes discover they need more than the template gave them, such as a second dataset or a public API. Handling this well keeps users from demanding broad templates up front. The agent can emit a capability request: a structured statement of what it wants and why. The request goes to the orchestrator, never to the sandbox, and is decided by rules first and a human second.
- Requests that the user's entitlements already cover, such as another dataset they can read, can be granted automatically by recompiling the policy and starting a fresh sandbox under the new hash.
- Requests outside entitlements, such as a new external host, go to an approver with the task, the justification and the content that led to the request. An injection that asks the agent to request access is then visible to a person.
- Grants never modify a running sandbox. A new execution starts under the new policy, so every execution maps to exactly one policy hash.
The output release gate
What leaves the sandbox is also a boundary. Output flows into the model's context, onto the user's screen or into storage, and each destination carries risk: oversized output inflates cost, crafted text can inject instructions into the next model call, and a file can smuggle data to a destination the network policy would never allow. The release gate checks declared outputs against the policy's types and sizes, truncates text with a visible marker, strips anything that is not a declared output, and labels content passed to the model as tool output rather than instructions.
Worked example: an injected spreadsheet
A finance user asks the agent to convert a supplier spreadsheet into euros and chart monthly totals. The task type is analytics; the policy compiles to analytics-readonly v7, with the one uploaded file as input and the internal FX-rate API as the only permitted host. A cell in the spreadsheet contains hidden text instructing the agent to send the file to an external URL.
| Step | What happens | Boundary that decides |
|---|---|---|
| 1 | The agent's code reads the file and fetches rates with GET; the proxy adds the credential | Proxy allow rule, credential broker |
| 2 | Following the injected text, the code tries a POST to an external host | Proxy: host-not-allowed, denial logged with rule id |
| 3 | The code tries to encode data into a GET query to the FX host | Proxy: request-too-large on URL length |
| 4 | The model, told of the denial, emits a capability request for the external host | Orchestrator: outside entitlements, sent to an approver with the source cell quoted |
| 5 | The chart and CSV are released; a stray archive in scratch is dropped | Release gate: types and declared outputs |
The task still succeeds, the attempted exfiltration fails at three separate points, and the audit log shows a cluster of denials under one task id, which is the signal that triggers a security review. Nothing in this outcome depended on the model recognising the injection.
Testing and versioning policies
A policy that has never been seen to deny something is an assumption. Keep a denial suite: small programs that attempt each forbidden action and must fail. Run it in CI on every change to a template, an image or the proxy, and on a schedule in production against a canary sandbox.
| Probe | Expected result |
|---|---|
| Open a TCP connection to an address outside the allowlist | Connection refused, denial logged |
| Resolve a name through a custom DNS server | No route; only the proxy resolves |
| Read environment variables and common credential paths | No secrets present |
| Write outside the scratch directory | Read-only file system error |
| Fork processes in a loop | Stopped at the pids limit |
| Follow a redirect from an allowed host to a disallowed one | Denied at the redirect |
| Exceed the egress byte budget with allowed requests | Denied with budget rule id |
Version every template, record the policy hash on every execution and every audit event, and roll template changes out gradually, watching denial rates per rule. A sudden rise in denials after a change usually means a template became too strict for a legitimate task; a sudden fall after a change deserves more suspicion.
Trade-offs and failure modes
- Broad templates to avoid friction. The most common failure. Measure how often capability requests are granted automatically, and split templates when one task type keeps widening.
- Policy chosen by the model. Any path where model output influences capabilities without passing entitlement checks is an injection route.
- Mounted host paths. Convenient and dangerous. Copy inputs in by reference instead.
- Network access for package installs. Bake dependencies into pinned images; treat runtime installs as a capability request.
- Unbounded outputs. Cost and injection risk. Cap sizes and label output as data.
- Cost of isolation. Strong isolation and fresh sandboxes add latency; pools of pre-started sandboxes help, but a pooled sandbox must never be reused across tasks or tenants.
The general guardrail picture, of which sandbox policy is one layer, is in agent guardrails. Trace each execution with its task id and policy hash, as described in agentic observability and tracing, so a denial can be followed back to the content that caused it.
What to do next
- Inventory where agent-written code runs today and what network, files and secrets it can reach.
- Write one declarative policy template per task type, with default-deny network and a pinned image.
- Compile policies in the orchestrator from task type, user entitlements and task inputs, and record the hash.
- Remove the sandbox's own network interface and route all egress through a per-task proxy.
- Move secrets into a broker that injects them at the proxy for allowed hosts only.
- Add a capability-request path with automatic grants inside entitlements and human approval outside them.
- Put a release gate on outputs, then build the denial suite and run it in CI and against a production canary.