Giving an ADK Java agent file tools is the easy part, and the details of doing it safely, with a workspace guard, bounded reads and atomic confirmed writes, are in the file tools article. That article ends on an honest caveat: the guard checks a path and then uses it in a separate system call, so it is not a sandbox. It also assumes one workspace directory on one machine. Production breaks both assumptions. Many sessions run at once, they resume on whichever replica the load balancer picks, files need to outlive the pod, and disks fill up.

This article designs the file system underneath the tools: what kinds of files an agent touches, a workspace per session keyed from the tool context, snapshots to the artifact service so a session can move between replicas, quotas, operating-system confinement that closes the gap the guard leaves, an MCP filesystem server as an out-of-process alternative, and cleanup. It reuses the reader's Workspace guard and file tools by reference.

Three kinds of files

Sort files by who writes them and how long they must live, because each kind wants a different home:

KindExamplesWritten byLifetimeHome
ReferencePolicy PDFs, product catalogue, code checkoutYour pipeline, never the agentUntil the next data releaseRead-only mount shared by all sessions
ScratchIntermediate CSVs, drafts, extracted textThe agent, through toolsOne sessionPer-session directory on local disk
DeliverableThe final report, a patch, an exportThe agent, then confirmedAs long as the user needs itArtifact service, versioned

Most incidents with agent files come from mixing these. An agent that can write into the reference directory can poison every later session. A deliverable left on local scratch disappears when the pod is rescheduled. Keeping the three apart is the main design decision; the rest is mechanics.

Architecture

Three kinds of files, three homes, one confined processagent container: non-root, read-only root file systemLlmAgent + Runnerfile toolsWorkspaceManagersession key -> directory/ref (ro)shared reference/work/{hash}per-session scratchquota ledgerbytes, filesreadread/writeMCP filesystem serverseparate process, own mountsArtifact serviceversioned snapshotsSession servicemanifest in stateSweeperTTL and orphan cleanupsnapshothydratepointerLocal disk is a cache; the artifact service is the record. The path guard runs inside a process that can see nothing else.
Reference data is read-only, scratch is per session and cached locally, deliverables and snapshots live in the artifact service.

Every file call goes through the manager, which picks the root from the session and the tool, never from the model's arguments. Local directories can vanish at any time without losing work, because the artifact service holds the latest snapshot.

A workspace per session

Derive the workspace from the session, never from anything the model says. ToolContext extends CallbackContext and, through ReadonlyContext, exposes userId() and sessionId(). Hash them into the directory name so that unusual characters in ids cannot become path syntax, and so directory listings on the host do not leak user ids.

public final class WorkspaceManager {
  private final Path scratchRoot;               // e.g. /work, an emptyDir with a size limit
  private final Path referenceRoot;             // e.g. /ref, mounted read-only
  private final Map<String, Workspace> open = new ConcurrentHashMap<>();

  public Workspace scratch(ToolContext ctx) {
    String key = ctx.userId() + "\u0000" + ctx.sessionId();
    return open.computeIfAbsent(key, k -> {
      try {
        Path dir = scratchRoot.resolve(sha256Hex(k).substring(0, 32));
        Files.createDirectories(dir, PosixFilePermissions.asFileAttribute(
            PosixFilePermissions.fromString("rwx------")));
        return new Workspace(dir);              // the guard from the file tools article
      } catch (IOException e) { throw new UncheckedIOException(e); }
    });
  }

  public Workspace reference() throws IOException { return new Workspace(referenceRoot); }
}

Give reads and writes separate tool instances: readReference(path, startLine, maxLines) resolves against the reference root, readScratch and writeScratch against the session's directory. Two roots in two tools are easier to reason about, and to audit, than one tool with a flag the model can set. The POSIX attribute call fails on Windows, so guard it if you develop there.

Snapshots so sessions can move

Local disk belongs to one pod. When the session's next turn lands on another replica, or the pod is replaced during a deploy, the scratch directory is gone. Treat local disk as a cache and the artifact service as the record. ADK's BaseArtifactService stores versioned binary Parts per app, user and session, and the tool context exposes saveArtifact(filename, part) returning Completable, loadArtifact(filename) returning Maybe<Part> and listArtifacts(). The artifacts article covers the service itself; here it backs the workspace.

static final String SNAPSHOT = "workspace.zip";

/** Called at the end of any tool that wrote to scratch. */
Completable snapshot(ToolContext ctx, Workspace ws) {
  return Completable.defer(() -> {
    byte[] zip = Zip.directory(ws.root(), MAX_SNAPSHOT_BYTES);       // refuses links, caps size
    return ctx.saveArtifact(SNAPSHOT, Part.fromBytes(zip, "application/zip"));
  }).subscribeOn(Schedulers.io());
}

/** Called before the first scratch access on this replica. */
Completable hydrate(ToolContext ctx, Workspace ws) {
  return ctx.loadArtifact(SNAPSHOT)
      .flatMapCompletable(part -> Completable.fromAction(() ->
          Zip.extractInto(ws.root(), part.inlineData().flatMap(Blob::data).orElseThrow(),
                          MAX_FILES, MAX_UNPACKED_BYTES)))               // zip-slip and zip-bomb checks
      .subscribeOn(Schedulers.io());
}

Two details matter. Extraction is the most dangerous code in this design: an archive entry named ../../etc/cron.d/x is a path traversal, and a small archive can expand to gigabytes, so run every entry name through the same Workspace.resolve guard and cap both file count and unpacked bytes. And FunctionTool unwraps Single and Maybe return values, so a write tool can return Single<Map<String, Object>> that completes only after the snapshot is saved; the model never sees success for a file that was not persisted.

Whole-directory snapshots are simple and fine for workspaces of a few megabytes. Larger workspaces want per-file artifacts plus a manifest in session state mapping file names to artifact versions, so a turn uploads only what changed.

Quotas

An agent in a loop can write until the disk is full, and a full shared disk takes down every session on the node. Enforce limits at two layers. In the tool, keep a per-session ledger and check it before writing; at the platform, give the scratch volume a hard size so a bug in the ledger cannot exhaust the node.

LimitTypical valueEnforced by
Bytes per file2 MBwriteScratch, before the temp file is opened
Files per session200Ledger count
Bytes per session50 MBLedger sum, recomputed on hydrate
Bytes written per turn10 MBLedger, reset at each invocation
Volume sizeSessions per pod times bytes per session, plus headroomKubernetes emptyDir sizeLimit or a tmpfs size

Return a quota error as a normal tool result with a code such as QUOTA_EXCEEDED and the remaining allowance, so the model can delete scratch files or summarise instead of retrying the same write. Per-tenant ceilings across sessions belong in the platform's quota system; see the quota management article.

Confinement the guard cannot provide

The guard's remaining hole is a race: between the check and the open, a symbolic link can replace a directory. Java can narrow it: java.nio.file.SecureDirectoryStream opens entries relative to an already-open directory and accepts NOFOLLOW_LINKS, but it is an optional feature that the default provider offers on some platforms (Linux, for example) and not others, such as Windows, so code must test for it with instanceof and fall back. Rather than depend on it, close the gap with the operating system, by making sure that even a perfect exploit of the guard reaches nothing worth having.

# Pod spec excerpt: the agent sees only its scratch volume and a read-only reference mount.
spec:
  containers:
    - name: agent
      securityContext:
        runAsNonRoot: true
        runAsUser: 10001
        readOnlyRootFilesystem: true
        allowPrivilegeEscalation: false
        capabilities: { drop: ["ALL"] }
      volumeMounts:
        - { name: work, mountPath: /work }
        - { name: ref,  mountPath: /ref, readOnly: true }
        - { name: tmp,  mountPath: /tmp }
  volumes:
    - { name: work, emptyDir: { sizeLimit: 4Gi } }
    - { name: tmp,  emptyDir: { sizeLimit: 256Mi } }
    - { name: ref,  persistentVolumeClaim: { claimName: reference-data, readOnly: true } }

With a read-only root, no secrets mounted as files in the same container, and the reference data mounted read-only, a traversal can at worst reach another session's scratch directory on the same pod. If that is unacceptable, because sessions belong to different tenants, run file operations in a separate process per session or per tenant, as the tenant isolation article describes for the runtime as a whole. The guard stays as the first layer; it produces clean errors for the common mistakes, and the operating system catches what it misses.

Using the MCP filesystem server

An alternative to writing file tools is the reference MCP filesystem server, which restricts operations to allowed directories and exposes tools including read_text_file, list_directory, search_files, get_file_info, write_file, edit_file and move_file. ADK Java's McpToolset accepts a tool-name list, which is the place to allowlist read-only tools.

StdioServerParameters fs = StdioServerParameters.builder()
    .command("npx")
    .args(List.of("-y", "@modelcontextprotocol/server-filesystem", "/ref"))
    .build();

McpToolset referenceFiles = new McpToolset(
    fs.toServerParameters(), JsonBaseModel.getMapper(),
    List.of("read_text_file", "list_directory", "search_files", "get_file_info"));

Three cautions. The server's README says that if the client supplies MCP roots, they replace the directories given on the command line entirely, so check whether your client sends roots before trusting the arguments. A stdio server shares your container's file system unless you run it in its own container with only the intended mounts, so the confinement section still applies. And the server's tool list can change when you upgrade it; pin its version. The MCP integration article covers transports and lifecycle.

Lifecycle and cleanup

Workspaces need an end. Delete the scratch directory when a session is deleted, and run a sweeper that removes directories untouched for longer than your session idle timeout, because most sessions are abandoned rather than closed. Because the directory name is a hash, the sweeper cannot map it back to a session; keep a small index file inside each workspace with the session key and last-access time, written by the manager, not the agent. Artifacts follow the artifact service's retention, which must match what you told users about how long files are kept.

Worked example: a report across a pod eviction

The sizes in this example are illustrative. A reporting agent answers "summarise last quarter's returns by region". Turn one reads returns_q3.csv (38 MB) from /ref through a windowed reader, writes by_region.csv (14 KB) to scratch, and the write tool snapshots a 6 KB zip as version 0 of workspace.zip. Between turns the pod is evicted. Turn two lands on another replica, which has no directory for the session; the first scratch access hydrates version 0, and the agent writes report.md, saved as version 1 and, after user confirmation, as a separate report.md deliverable artifact.

In turn three the user asks for one chart per country. The model loops, writing a 900 KB image per call; after 11 images the turn has written 9.9 MB, so the 12th call would pass the 10 MB per-turn allowance and the tool returns QUOTA_EXCEEDED with 0.1 MB left for this turn. The model reports the limit and offers charts for the top 10 countries instead, with the rest in a table. The node's other sessions never notice, because the volume limit was never approached.

Failure modes

FailureSymptomGuard
Agent writes to reference dataEvery later session sees poisoned filesRead-only mount; separate reference tools
Session moves replicasFiles missing on the next turnSnapshot on write, hydrate on first access
Zip-slip or zip bomb on hydrateFiles outside the workspace or a full diskGuard every entry; cap count and bytes
Runaway writesNode disk full, all sessions failPer-session ledger plus volume sizeLimit
Check-then-use raceRead outside the workspaceNon-root, read-only root, minimal mounts
MCP roots overrideServer reaches unexpected directoriesCheck roots behaviour; isolate the server's mounts
Abandoned workspacesDisk slowly fillsIndex file and TTL sweeper

Trade-offs

Snapshotting after every write adds latency and artifact storage; snapshotting at the end of a turn is cheaper but loses a turn's work if the pod dies mid-turn. Whole-directory zips are simple but scale poorly past tens of megabytes. Writing your own tools gives you exact verbs and error codes, while the MCP server gives you a maintained implementation with a broader surface that you must narrow. A network file system shared by all replicas removes hydration entirely, but reintroduces cross-session exposure and adds latency to every read; it suits large reference data better than scratch.

What to do next

  1. Classify every file your agent touches as reference, scratch or deliverable, and give each its own root.
  2. Key scratch directories from a hash of user and session id taken from the tool context.
  3. Snapshot scratch to the artifact service on write, hydrate on first access, and run every archive entry through the guard.
  4. Add a per-session quota ledger and a hard size limit on the scratch volume.
  5. Run the agent non-root with a read-only root file system and a read-only reference mount.
  6. If you use the MCP filesystem server, allowlist read-only tools, pin its version and check its roots behaviour.
  7. Write the sweeper, then test a replica failover and a quota breach end to end.
Key takeaway: An agent's file system is three stores, not one: read-only reference data, per-session scratch and versioned deliverables. Key scratch from the session, treat local disk as a cache backed by artifact snapshots, enforce quotas in the tool and at the volume, and confine the process so the path guard is a first layer rather than the only one. Then sessions survive replica moves, runaway loops hit a clear limit, and a traversal bug reaches nothing worth having.