Giving an ADK Java agent file tools is the easy part, and the details of doing it safely, with a workspace guard, bounded reads and atomic confirmed writes, are in the file tools article. That article ends on an honest caveat: the guard checks a path and then uses it in a separate system call, so it is not a sandbox. It also assumes one workspace directory on one machine. Production breaks both assumptions. Many sessions run at once, they resume on whichever replica the load balancer picks, files need to outlive the pod, and disks fill up.
This article designs the file system underneath the tools: what kinds of files an agent touches, a workspace per session keyed from the tool context, snapshots to the artifact service so a session can move between replicas, quotas, operating-system confinement that closes the gap the guard leaves, an MCP filesystem server as an out-of-process alternative, and cleanup. It reuses the reader's Workspace guard and file tools by reference.
Three kinds of files
Sort files by who writes them and how long they must live, because each kind wants a different home:
| Kind | Examples | Written by | Lifetime | Home |
|---|---|---|---|---|
| Reference | Policy PDFs, product catalogue, code checkout | Your pipeline, never the agent | Until the next data release | Read-only mount shared by all sessions |
| Scratch | Intermediate CSVs, drafts, extracted text | The agent, through tools | One session | Per-session directory on local disk |
| Deliverable | The final report, a patch, an export | The agent, then confirmed | As long as the user needs it | Artifact service, versioned |
Most incidents with agent files come from mixing these. An agent that can write into the reference directory can poison every later session. A deliverable left on local scratch disappears when the pod is rescheduled. Keeping the three apart is the main design decision; the rest is mechanics.
Architecture
Every file call goes through the manager, which picks the root from the session and the tool, never from the model's arguments. Local directories can vanish at any time without losing work, because the artifact service holds the latest snapshot.
A workspace per session
Derive the workspace from the session, never from anything the model says. ToolContext extends CallbackContext and, through ReadonlyContext, exposes userId() and sessionId(). Hash them into the directory name so that unusual characters in ids cannot become path syntax, and so directory listings on the host do not leak user ids.
public final class WorkspaceManager {
private final Path scratchRoot; // e.g. /work, an emptyDir with a size limit
private final Path referenceRoot; // e.g. /ref, mounted read-only
private final Map<String, Workspace> open = new ConcurrentHashMap<>();
public Workspace scratch(ToolContext ctx) {
String key = ctx.userId() + "\u0000" + ctx.sessionId();
return open.computeIfAbsent(key, k -> {
try {
Path dir = scratchRoot.resolve(sha256Hex(k).substring(0, 32));
Files.createDirectories(dir, PosixFilePermissions.asFileAttribute(
PosixFilePermissions.fromString("rwx------")));
return new Workspace(dir); // the guard from the file tools article
} catch (IOException e) { throw new UncheckedIOException(e); }
});
}
public Workspace reference() throws IOException { return new Workspace(referenceRoot); }
}Give reads and writes separate tool instances: readReference(path, startLine, maxLines) resolves against the reference root, readScratch and writeScratch against the session's directory. Two roots in two tools are easier to reason about, and to audit, than one tool with a flag the model can set. The POSIX attribute call fails on Windows, so guard it if you develop there.
Snapshots so sessions can move
Local disk belongs to one pod. When the session's next turn lands on another replica, or the pod is replaced during a deploy, the scratch directory is gone. Treat local disk as a cache and the artifact service as the record. ADK's BaseArtifactService stores versioned binary Parts per app, user and session, and the tool context exposes saveArtifact(filename, part) returning Completable, loadArtifact(filename) returning Maybe<Part> and listArtifacts(). The artifacts article covers the service itself; here it backs the workspace.
static final String SNAPSHOT = "workspace.zip";
/** Called at the end of any tool that wrote to scratch. */
Completable snapshot(ToolContext ctx, Workspace ws) {
return Completable.defer(() -> {
byte[] zip = Zip.directory(ws.root(), MAX_SNAPSHOT_BYTES); // refuses links, caps size
return ctx.saveArtifact(SNAPSHOT, Part.fromBytes(zip, "application/zip"));
}).subscribeOn(Schedulers.io());
}
/** Called before the first scratch access on this replica. */
Completable hydrate(ToolContext ctx, Workspace ws) {
return ctx.loadArtifact(SNAPSHOT)
.flatMapCompletable(part -> Completable.fromAction(() ->
Zip.extractInto(ws.root(), part.inlineData().flatMap(Blob::data).orElseThrow(),
MAX_FILES, MAX_UNPACKED_BYTES))) // zip-slip and zip-bomb checks
.subscribeOn(Schedulers.io());
}Two details matter. Extraction is the most dangerous code in this design: an archive entry named ../../etc/cron.d/x is a path traversal, and a small archive can expand to gigabytes, so run every entry name through the same Workspace.resolve guard and cap both file count and unpacked bytes. And FunctionTool unwraps Single and Maybe return values, so a write tool can return Single<Map<String, Object>> that completes only after the snapshot is saved; the model never sees success for a file that was not persisted.
Whole-directory snapshots are simple and fine for workspaces of a few megabytes. Larger workspaces want per-file artifacts plus a manifest in session state mapping file names to artifact versions, so a turn uploads only what changed.
Quotas
An agent in a loop can write until the disk is full, and a full shared disk takes down every session on the node. Enforce limits at two layers. In the tool, keep a per-session ledger and check it before writing; at the platform, give the scratch volume a hard size so a bug in the ledger cannot exhaust the node.
| Limit | Typical value | Enforced by |
|---|---|---|
| Bytes per file | 2 MB | writeScratch, before the temp file is opened |
| Files per session | 200 | Ledger count |
| Bytes per session | 50 MB | Ledger sum, recomputed on hydrate |
| Bytes written per turn | 10 MB | Ledger, reset at each invocation |
| Volume size | Sessions per pod times bytes per session, plus headroom | Kubernetes emptyDir sizeLimit or a tmpfs size |
Return a quota error as a normal tool result with a code such as QUOTA_EXCEEDED and the remaining allowance, so the model can delete scratch files or summarise instead of retrying the same write. Per-tenant ceilings across sessions belong in the platform's quota system; see the quota management article.
Confinement the guard cannot provide
The guard's remaining hole is a race: between the check and the open, a symbolic link can replace a directory. Java can narrow it: java.nio.file.SecureDirectoryStream opens entries relative to an already-open directory and accepts NOFOLLOW_LINKS, but it is an optional feature that the default provider offers on some platforms (Linux, for example) and not others, such as Windows, so code must test for it with instanceof and fall back. Rather than depend on it, close the gap with the operating system, by making sure that even a perfect exploit of the guard reaches nothing worth having.
# Pod spec excerpt: the agent sees only its scratch volume and a read-only reference mount.
spec:
containers:
- name: agent
securityContext:
runAsNonRoot: true
runAsUser: 10001
readOnlyRootFilesystem: true
allowPrivilegeEscalation: false
capabilities: { drop: ["ALL"] }
volumeMounts:
- { name: work, mountPath: /work }
- { name: ref, mountPath: /ref, readOnly: true }
- { name: tmp, mountPath: /tmp }
volumes:
- { name: work, emptyDir: { sizeLimit: 4Gi } }
- { name: tmp, emptyDir: { sizeLimit: 256Mi } }
- { name: ref, persistentVolumeClaim: { claimName: reference-data, readOnly: true } }With a read-only root, no secrets mounted as files in the same container, and the reference data mounted read-only, a traversal can at worst reach another session's scratch directory on the same pod. If that is unacceptable, because sessions belong to different tenants, run file operations in a separate process per session or per tenant, as the tenant isolation article describes for the runtime as a whole. The guard stays as the first layer; it produces clean errors for the common mistakes, and the operating system catches what it misses.
Using the MCP filesystem server
An alternative to writing file tools is the reference MCP filesystem server, which restricts operations to allowed directories and exposes tools including read_text_file, list_directory, search_files, get_file_info, write_file, edit_file and move_file. ADK Java's McpToolset accepts a tool-name list, which is the place to allowlist read-only tools.
StdioServerParameters fs = StdioServerParameters.builder()
.command("npx")
.args(List.of("-y", "@modelcontextprotocol/server-filesystem", "/ref"))
.build();
McpToolset referenceFiles = new McpToolset(
fs.toServerParameters(), JsonBaseModel.getMapper(),
List.of("read_text_file", "list_directory", "search_files", "get_file_info"));Three cautions. The server's README says that if the client supplies MCP roots, they replace the directories given on the command line entirely, so check whether your client sends roots before trusting the arguments. A stdio server shares your container's file system unless you run it in its own container with only the intended mounts, so the confinement section still applies. And the server's tool list can change when you upgrade it; pin its version. The MCP integration article covers transports and lifecycle.
Lifecycle and cleanup
Workspaces need an end. Delete the scratch directory when a session is deleted, and run a sweeper that removes directories untouched for longer than your session idle timeout, because most sessions are abandoned rather than closed. Because the directory name is a hash, the sweeper cannot map it back to a session; keep a small index file inside each workspace with the session key and last-access time, written by the manager, not the agent. Artifacts follow the artifact service's retention, which must match what you told users about how long files are kept.
Worked example: a report across a pod eviction
The sizes in this example are illustrative. A reporting agent answers "summarise last quarter's returns by region". Turn one reads returns_q3.csv (38 MB) from /ref through a windowed reader, writes by_region.csv (14 KB) to scratch, and the write tool snapshots a 6 KB zip as version 0 of workspace.zip. Between turns the pod is evicted. Turn two lands on another replica, which has no directory for the session; the first scratch access hydrates version 0, and the agent writes report.md, saved as version 1 and, after user confirmation, as a separate report.md deliverable artifact.
In turn three the user asks for one chart per country. The model loops, writing a 900 KB image per call; after 11 images the turn has written 9.9 MB, so the 12th call would pass the 10 MB per-turn allowance and the tool returns QUOTA_EXCEEDED with 0.1 MB left for this turn. The model reports the limit and offers charts for the top 10 countries instead, with the rest in a table. The node's other sessions never notice, because the volume limit was never approached.
Failure modes
| Failure | Symptom | Guard |
|---|---|---|
| Agent writes to reference data | Every later session sees poisoned files | Read-only mount; separate reference tools |
| Session moves replicas | Files missing on the next turn | Snapshot on write, hydrate on first access |
| Zip-slip or zip bomb on hydrate | Files outside the workspace or a full disk | Guard every entry; cap count and bytes |
| Runaway writes | Node disk full, all sessions fail | Per-session ledger plus volume sizeLimit |
| Check-then-use race | Read outside the workspace | Non-root, read-only root, minimal mounts |
| MCP roots override | Server reaches unexpected directories | Check roots behaviour; isolate the server's mounts |
| Abandoned workspaces | Disk slowly fills | Index file and TTL sweeper |
Trade-offs
Snapshotting after every write adds latency and artifact storage; snapshotting at the end of a turn is cheaper but loses a turn's work if the pod dies mid-turn. Whole-directory zips are simple but scale poorly past tens of megabytes. Writing your own tools gives you exact verbs and error codes, while the MCP server gives you a maintained implementation with a broader surface that you must narrow. A network file system shared by all replicas removes hydration entirely, but reintroduces cross-session exposure and adds latency to every read; it suits large reference data better than scratch.
What to do next
- Classify every file your agent touches as reference, scratch or deliverable, and give each its own root.
- Key scratch directories from a hash of user and session id taken from the tool context.
- Snapshot scratch to the artifact service on write, hydrate on first access, and run every archive entry through the guard.
- Add a per-session quota ledger and a hard size limit on the scratch volume.
- Run the agent non-root with a read-only root file system and a read-only reference mount.
- If you use the MCP filesystem server, allowlist read-only tools, pin its version and check its roots behaviour.
- Write the sweeper, then test a replica failover and a quota breach end to end.