An agent with a run_code tool executes text written by a model, and the model reads content an attacker can influence: web pages, documents, tool results. Treat every snippet it runs as hostile. The sandbox is the only thing between that code and your host, your credentials and your network.

This page builds that sandbox on a single host with Docker and gVisor, one control at a time, and shows how to prove each control works. It is deliberately concrete. The wider architecture, with policy gates, microVM pools and brokered egress, is in Agent tool-execution sandboxing architecture, and deriving per-task policy is in Sandboxing Code-Executing Agents with Policy Boundaries. If you run on Kubernetes, the same controls appear as pod settings in Container security architecture.

Advertisement

The threat model for a run_code tool

Assume the attacker fully controls the code. What they want, in rough order of damage:

  1. Escape to the host through a kernel bug, a privileged capability, a mounted Docker socket or a writable host path.
  2. Steal secrets from environment variables, mounted credential files, the image itself, or a cloud metadata endpoint such as 169.254.169.254.
  3. Use the network to exfiltrate data, reach internal services, or download a second stage.
  4. Exhaust resources: fork bombs, memory, disk, CPU, or an output flood that fills the runner's memory or the model's context.
  5. Persist into a later run, for example by writing to a container or volume that the next user's code shares.
  6. Manipulate the agent by printing text that reads like instructions. Sandbox output is untrusted input to the next model call.

A default docker run stops few of these. The process runs as root inside the container, keeps a default set of capabilities, has a bridge network with full outbound access, a writable root filesystem and no process limit. Most importantly, containers share the host kernel, so a kernel vulnerability reachable through an allowed syscall is a host compromise. Docker's default seccomp profile blocks a set of dangerous syscalls, but hundreds remain available.

Why gVisor, and what it costs

gVisor's runsc runtime puts a user-space kernel, the Sentry, between the container and the host. The sandboxed process's syscalls are intercepted and handled by the Sentry, which itself uses a small, filtered set of host syscalls. A bug in Linux's implementation of some rarely used syscall is no longer directly reachable from attacker code. The default interception platform is systrap, which uses seccomp traps; a KVM platform also exists and is configured with a runtime flag.

The costs are real. Syscall-heavy workloads (lots of small file operations, many processes) run noticeably slower, and some kernel features and syscalls are not implemented, so a few programs fail under runsc that run fine under runc. For short data-analysis scripts this is usually acceptable. If you need stronger isolation with a separate guest kernel per run, microVMs such as Firecracker or Kata Containers are the next step up, at higher operational cost.

# Install runsc from the gVisor release channel (see gvisor.dev), then:
sudo runsc install                 # adds a "runsc" runtime to /etc/docker/daemon.json
sudo systemctl restart docker

# Resulting daemon.json entry
{
  "runtimes": {
    "runsc": { "path": "/usr/local/bin/runsc" }
  }
}

# Verify: under gVisor, dmesg shows gVisor's own boot messages, not the host's
docker run --rm --runtime=runsc alpine dmesg
One run_code call: where each control sitsAgent loopmodel emits codeRunner (trusted)timeout, output cap, rm -fcodeDocker daemonflags become configdocker runSandbox container (one per call, removed after)python3 -I -uid 65534, no capsread-only roottmpfs /tmp, noexec, 64 MBnetwork=noneno route, no DNSpids 128, 512 MB, 1 CPUcgroup limitsrunsc (gVisor Sentry)answers most syscalls in user spaceHost kernelsyscallsstdout/stderr (capped)Output goes back to the modelas untrusted data
Every control is either a docker run flag, a runner behaviour, or the runsc boundary underneath.
Advertisement

The run command, flag by flag

Each flag closes one avenue. The third column is the probe you run inside the sandbox to prove the flag took effect; the next sections turn those probes into a test.

FlagWhat it blocksProbe
--runtime=runscDirect host-kernel syscall attack surfacedmesg shows gVisor messages
--network=noneExfiltration, internal scanning, downloads, metadata endpointTCP connect to any address fails
--cap-drop=ALLRaw sockets, chown, mknod and other privileged operationsCapEff in /proc/self/status is zero
--security-opt=no-new-privilegesGaining privilege through setuid binariesNoNewPrivs: 1 where the kernel view exposes it
--user 65534:65534Root inside the containeros.getuid() is 65534
--read-onlyModifying the image, planting files for later runsthe / entry in /proc/self/mounts has ro
--tmpfs /tmp:rw,noexec,nosuid,size=64mUnbounded scratch and executing dropped binarieswriting 65 MB fails; exec from /tmp fails
--pids-limit=128Fork bombsforking 200 children fails with EAGAIN
--memory=512m --memory-swap=512mMemory exhaustion of the hostallocating 1 GB kills the process, exit 137
--cpus=1CPU starvation of neighboursbusy loop shows one core in docker stats
--ulimit nofile=256:256Descriptor exhaustionopening 300 files fails
--ulimit fsize=16777216Single huge fileswriting a 20 MB file fails
--rm plus --namePersistence and leaked containersdocker ps -a is empty after each run

Equally important is what is absent: no -v /var/run/docker.sock (the socket is root on the host), no --privileged, no host paths, no -e secrets, no --network=host, and no --pid=host. Build the image from a minimal base, run as a non-root user, include only the libraries the agent is allowed to use, and pin it by digest so a registry change cannot swap it. On the daemon side, run Docker rootless or enable user-namespace remapping so that a container breakout lands as an unprivileged host user; check that your Docker version supports the combination with runsc before relying on it.

A runner that really stops the container

The most common bug in sandbox runners is the timeout. subprocess.run(["docker", "run", ...], timeout=10) kills the docker CLI client when time expires. The container keeps running on the daemon, burning CPU and holding memory, and an attacker who loops forever accumulates containers. Give every container a unique name and remove it by name in a finally block. The second common bug is capture_output=True, which buffers unlimited output in the runner's memory. Read at most a cap and close the pipe.

import subprocess, threading, uuid

IMAGE = "registry.example.com/sbx-python@sha256:<digest>"   # pin by digest

def sandbox_argv(name):
    return [
        "docker", "run", "--rm", "-i", "--name", name,
        "--runtime=runsc", "--network=none",
        "--cap-drop=ALL", "--security-opt=no-new-privileges",
        "--read-only", "--tmpfs", "/tmp:rw,noexec,nosuid,size=64m",
        "--user", "65534:65534",
        "--pids-limit=128", "--memory=512m", "--memory-swap=512m", "--cpus=1",
        "--ulimit", "nofile=256:256", "--ulimit", "fsize=16777216",
        IMAGE, "python3", "-I", "-",          # code arrives on stdin, never on argv
    ]

def _drain(stream, cap, out):
    out.append(stream.read(cap + 1))         # stop reading after cap + 1 bytes
    stream.close()                           # writer then gets a broken pipe

def run_untrusted(code, timeout_s=10, cap=64_000):
    name = "sbx-" + uuid.uuid4().hex[:12]
    proc = subprocess.Popen(sandbox_argv(name), stdin=subprocess.PIPE,
                            stdout=subprocess.PIPE, stderr=subprocess.PIPE)
    out, err = [], []
    readers = [threading.Thread(target=_drain, args=(proc.stdout, cap, out)),
               threading.Thread(target=_drain, args=(proc.stderr, cap, err))]
    for t in readers:
        t.start()
    status = "ok"
    try:
        proc.stdin.write(code.encode())
        proc.stdin.close()
        proc.wait(timeout=timeout_s)
    except subprocess.TimeoutExpired:
        status = "timeout"
    except BrokenPipeError:
        status = "closed_early"
    finally:
        # Killing the CLI does not stop the container: remove it by name.
        subprocess.run(["docker", "rm", "-f", name], capture_output=True)
        proc.kill()
        proc.wait()
        for t in readers:
            t.join()
    if status == "ok" and (len(out[0]) > cap or len(err[0]) > cap):
        status = "output_truncated"
    return {"status": status, "exit": proc.returncode,
            "stdout": out[0][:cap].decode(errors="replace"),
            "stderr": err[0][:cap].decode(errors="replace")}

The process logic was tested with a local interpreter substituted for the docker command: normal output, a sleep past the timeout, and an infinite print loop that is cut off at the cap. Exit code 137 means the process received SIGKILL, from the memory limit or from your removal; record which by checking whether the status is timeout. Pass code on stdin rather than as an argument so it never appears in process listings or shell parsing. Return the result to the agent wrapped as data, labelled as program output, and never let the runner itself interpret it.

Probes that prove the boundary

Configuration drifts: someone adds a volume for debugging, a base image update brings a setuid binary, a Docker upgrade changes a default. Run a probe script through the real runner in CI and on a schedule in production, and fail loudly if any control is missing.

import json

PROBE = r"""
import json, os, signal, socket, time
results = {}
results["uid"] = os.getuid()
root = [l.split() for l in open("/proc/self/mounts") if l.split()[1] == "/"]
results["root_ro"] = "ro" in root[-1][3].split(",") if root else False
try:
    socket.create_connection(("1.1.1.1", 53), timeout=2); results["net"] = True
except OSError:
    results["net"] = False
caps = [l for l in open("/proc/self/status") if l.startswith("CapEff")]
results["capeff"] = caps[0].split()[1] if caps else "missing"
kids = []
try:
    for _ in range(200):
        pid = os.fork()
        if pid == 0:
            time.sleep(5)
            os._exit(0)
        kids.append(pid)
    results["pids_limited"] = False
except BlockingIOError:                      # EAGAIN only; other errors crash the probe
    results["pids_limited"] = True
finally:
    for pid in kids:
        os.kill(pid, signal.SIGKILL)
        os.waitpid(pid, 0)
print(json.dumps(results))
"""

def check_sandbox():
    r = run_untrusted(PROBE, timeout_s=20)
    got = json.loads(r["stdout"])
    assert got["uid"] == 65534
    assert got["root_ro"] is True
    assert got["net"] is False
    assert got["capeff"] == "0000000000000000"       # "missing" fails closed
    assert got["pids_limited"] is True

Run the probe under both runc and runsc when you first adopt gVisor. gVisor emulates /proc, so some fields can differ from a native kernel; the asserts are written so a missing field fails rather than passes. The probe emits JSON and the checker parses it with json.loads, never eval: even probe output passes through the sandbox and is untrusted.

Files and network when the task needs them

Input files. Avoid bind-mounting shared host directories. Create a fresh per-run directory, copy in only the files the task needs, mount it with -v /srv/sbx/<run>/in:/in:ro, and delete it afterwards. Never reuse a directory across users or sessions.

Output files. Have the code write to /tmp/out and print a tar stream or base64 to stdout under the same cap, or mount a second per-run directory read-write with a size-limited filesystem. Treat every returned file as untrusted: check type and size, and never render HTML or execute anything it contains on the host.

Network. Keep --network=none as the default and pre-install packages into the image instead of allowing pip install. If a task truly needs HTTP, attach the container to an internal Docker network whose only reachable host is an egress proxy enforcing an allowlist, as described in LLM egress filtering architecture. Block the metadata address at the proxy and on the host firewall, and never give the sandbox DNS that resolves internal names.

Operating it

  • Log every run with a hash of the code, the requesting session, status, exit code, duration and output size. Timeouts and 137 exits clustered on one session are a signal worth alerting on.
  • Bound concurrency per tenant and globally. Each sandbox reserves up to 512 MB here; ten parallel agents can claim 5 GB.
  • Measure start latency. Container start under runsc adds latency per call. Keep images small and pre-pulled. If you build a warm pool, hand each container to exactly one run and destroy it afterwards; never return a used container to the pool.
  • Patch both layers. Track runsc releases and host kernel updates; gVisor reduces kernel exposure but does not remove it.
  • Rehearse a kill. The runner should stop accepting work when a control fails its probe, and operators should be able to stop all sandboxes at once, as covered in Agent Kill Switch.

Failure modes and trade-offs

  • Orphaned containers from client-only timeouts. Fix: name and remove.
  • Runner out of memory from unbounded output capture. Fix: capped reads.
  • A debug volume left in production. Fix: probes in CI and on a schedule.
  • Secrets in the image or environment. Fix: build images with no credentials and give tools that need credentials a broker outside the sandbox.
  • Programs that fail under runsc because a syscall is unimplemented. Fix: test your library set under runsc, and do not silently fall back to runc.
  • Prompt injection through output. Fix: wrap output as data and keep the agent's other tools on least privilege, as in Agent Tool Permissions.
OptionBoundaryOverheadGood for
runc plus the flags aboveNamespaces, cgroups, seccomp; shared kernelLowestTrusted code, or a first step while adopting gVisor
runsc plus the flags aboveUser-space kernel in front of the host kernelModerate on syscall-heavy codeModel-written scripts on a single host
MicroVM (Firecracker, Kata)Separate guest kernel per runHighest operational costMulti-tenant platforms and hostile workloads at scale

What to do next

  1. Install runsc on a test host and confirm dmesg inside a container shows gVisor.
  2. Copy the flag set above and remove anything your tasks do not need, starting from no network.
  3. Build a minimal, non-root image pinned by digest with only the allowed libraries.
  4. Replace any subprocess.run(..., timeout=) runner with one that names and removes containers and caps output.
  5. Write the probe script, emit JSON, and run it in CI and every hour in production.
  6. Decide how files enter and leave, with per-run directories deleted after each call.
  7. Add logs and alerts for timeouts, 137 exits, output truncation and concurrency.
  8. Review the wider architecture and decide whether microVMs are needed for your tenancy model.
Key takeaway: Treat model-written code as hostile and give it one disposable container per call: gVisor in front of the host kernel, no network, no capabilities, no new privileges, a non-root user, a read-only root with a small noexec tmpfs, and limits on processes, memory, CPU, files and output. Make the runner remove containers by name on timeout and cap what it reads, prove every control with a probe run through the real runner, and keep secrets and network access outside the sandbox.