A code execution tool lets the model answer by writing a short program and reading its output, instead of guessing at arithmetic or inventing a table. Ask an agent for the monthly payment on a loan, the median of three hundred numbers or the date ninety business days from now, and a language model that computes in its head will often be close and sometimes badly wrong. A model that writes ten lines of Python and runs them is exact, and the code it ran is a record you can audit.
This page explains how ADK Java implements that loop in version 1.11.0, which executor to pick, how the per-invocation error budget behaves, where output files go and what fails in production. Every class name and default below was checked against the 1.11.0 jar with javap. Isolation, meaning what stops model-written code from reading your secrets, has its own page: tool sandboxing in ADK Java. Read it before you run generated code anywhere near production.
Code executors versus function tools
A code executor is not a FunctionTool. A function tool is code you wrote; the model only chooses arguments. A code executor runs code the model wrote. In ADK Java you attach it to an agent with LlmAgent.builder().codeExecutor(...), and it takes part in the request and response processing of every model call that agent makes.
Use it when the task is computation over data the model already has: statistics, unit conversion, date arithmetic, reshaping a small table, checking a formula. Do not use it as a general integration mechanism. If the agent needs to look up an order, write a function tool with a schema. Generated code that opens network connections is exactly what a good sandbox forbids, so an agent that needs both should have both: function tools for systems, a code executor for maths.
The execution loop, step by step
The work happens in the CodeExecution request and response processors in com.google.adk.flows.llmflows. For every executor except the built-in one, one turn goes like this:
- The model returns a response. The post-processor looks for the first code block, using the executor's
codeBlockDelimiters(). The defaults are fenced blocks opening with```tool_codeor```python. - If there is a block, the response content is cut off after it and emitted as an event authored by the agent. Anything the model wrote after the code, such as an invented result, is discarded. That is deliberate: the real output comes next.
- The executor's
executeCode(InvocationContext, CodeExecutionInput)runs the code and returns aCodeExecutionResultwithstdout(),stderr()andoutputFiles(). - ADK emits a second event with role
modelcarrying acodeExecutionResultpart, plus a state delta recording the execution. Output files are saved through the artifact service. - The flow calls the model again with the code and its result in history, and the model either writes more code or answers.
If the response contains no code block, nothing happens and the turn ends normally. The request processor has a second job, preparing inline CSV files from the user message, but only when optimizeDataFile() is true. It is false for all three shipped executors, so in practice that path only runs for a custom executor that turns it on.
Choosing an executor
ADK Java 1.11.0 ships three executors, and you can write a fourth. They differ in where the code runs and in what comes back.
| Executor | Where code runs | What to know |
|---|---|---|
BuiltInCodeExecutor | On Google's side, inside the Gemini call | Adds Gemini's code execution tool to the request. Its own executeCode throws UnsupportedOperationException; Gemini returns the code and result parts itself. Gemini models only; it rejects other models with IllegalArgumentException. |
ContainerCodeExecutor | A Docker container you control | Runs python3 -c <code> through docker-java. Not stateful. Per the 1.11.0 bytecode it never reads input files or returns output files. |
VertexAiCodeExecutor | Vertex AI Code Interpreter extension | Takes the extension resource name, or reads CODE_INTERPRETER_EXTENSION_NAME. Returns output files, recognising png, jpg and jpeg images and csv data. Uses the v1beta1 extensions API. |
Your subclass of BaseCodeExecutor | Anywhere | Implement one method. Useful for wrapping, filtering and metrics. |
There is also BuiltInCodeExecutionTool, a tool rather than an executor, which adds the same Gemini code execution capability through the tool list and logs an error for non-Gemini models. Use one Gemini route, not both. The built-in route is the simplest and keeps code off your infrastructure, but you get no hook on the code before it runs and the environment is whatever Gemini provides. The container route is the one to choose when you need a fixed library set, your own audit log, or a model other than Gemini.
Wiring it into an agent
Wiring is two lines on the agent and one requirement on the runner. Runner.Builder refuses to build without an artifact service ("Artifact service must be provided."), and the code execution path needs one because output files are saved as artifacts. InMemoryRunner supplies an in-memory one, which is fine for tests and loses files on restart.
ContainerCodeExecutor executor = ContainerCodeExecutor
.fromImage("registry.example.com/agent-python:3.12-2026-10")
.setStrictSandbox(true) // one locked-down container per run
.setExecutionTimeoutSeconds(20) // strict mode only
.setMemoryLimitBytes(256L * 1024 * 1024);
LlmAgent analyst = LlmAgent.builder()
.name("loan_analyst")
.model("gemini-2.5-flash")
.instruction("""
You answer questions about loans. For any arithmetic, write Python in a
```python block that prints the answer. Never state a computed number
you did not print. Use only the standard library.
""")
.codeExecutor(executor)
.build();
Runner runner = Runner.builder()
.appName("loans")
.agent(analyst)
.sessionService(new InMemorySessionService())
.artifactService(new InMemoryArtifactService())
.build();
// executor implements AutoCloseable: close it on shutdown to remove containersTwo instruction lines matter more than they look. "Print the answer" matters because the executor only captures stdout; a bare expression on the last line prints nothing under python3 -c. "Never state a number you did not print" ties the final answer to the execution record. Pin the image by tag or digest so the library set your prompt promises is the one that exists.
ContainerCodeExecutor in practice
ContainerCodeExecutor has two modes, and the default is the unsafe one. With strict mode off, which is the 1.11.0 default, every execution runs in one shared container with network access, a writable filesystem and no memory, process or time limits. The executor logs a warning saying so, and says strict mode becomes the default in ADK 2.0.
With setStrictSandbox(true) each execution gets its own container with all capabilities dropped, a read-only root filesystem, no-new-privileges, a 128-process limit, a 64 MB writable /tmp, networking off unless you call setNetworkEnabled(true), a memory limit (512 MiB by default) and an execution timeout (60 seconds by default). The timeout and memory setters do nothing in shared mode. The exact host configuration, and why the JVM cannot sandbox this itself, are covered in the sandboxing guide.
Because stateful() is false, variables do not survive between runs: each block must be self-contained. Because the executor ignores input and output files, data has to travel inside the code, and charts cannot come back. If you need either, use the Vertex executor or wrap the container one.
The error budget
The post-processor keeps an error count per invocation in session state. Any non-empty stderr increments it; a clean run resets it. Once the count reaches errorRetryAttempts(), which is 2 by default, the post-processor stops running code for the rest of that invocation. The model can still write code blocks; they are just not executed, so the user may see Python where an answer should be.
The trap is the word "any". A library that prints a deprecation warning to stderr counts as an error, so two warnings in a row disable execution even though both runs succeeded. The cure is not a bigger retry number. Silence known warnings in the image, for example with PYTHONWARNINGS=ignore::DeprecationWarning in its environment, and classify stderr in a wrapper so only real failures count. The wrapper below does that, and also caps output so a loop that prints a million lines cannot fill the context window.
A guarded custom executor
public final class GuardedExecutor extends BaseCodeExecutor {
private static final int MAX_CHARS = 4_000;
private static final Pattern REAL_ERROR =
Pattern.compile("Traceback|Error:|timed out|interrupted");
private final BaseCodeExecutor delegate;
public GuardedExecutor(BaseCodeExecutor delegate) { this.delegate = delegate; }
@Override
public CodeExecutionResult executeCode(InvocationContext ctx, CodeExecutionInput input) {
long start = System.nanoTime();
CodeExecutionResult r = delegate.executeCode(ctx, input);
String err = r.stderr() == null ? "" : r.stderr();
// warnings alone must not spend the error budget
String kept = REAL_ERROR.matcher(err).find() ? clip(err) : "";
AuditLog.record(ctx.invocationId(), input.code(), r.stdout(), err,
Duration.ofNanos(System.nanoTime() - start)); // your own sink
return CodeExecutionResult.builder()
.stdout(clip(r.stdout()))
.stderr(kept)
.outputFiles(r.outputFiles())
.build();
}
@Override public int errorRetryAttempts() { return 3; }
private static String clip(String s) {
if (s == null || s.length() <= MAX_CHARS) return s == null ? "" : s;
return s.substring(0, MAX_CHARS) + "\n[output truncated at " + MAX_CHARS + " chars]";
}
}The audit record is the important part. It stores the exact code the model ran next to the answer it gave, which is what you need when a user disputes a number. Keep the pattern conservative: it is better to count a harmless line as an error than to hide a real traceback from the model, which uses stderr to fix its own code.
Worked example: a loan payment
A user asks: "What is the monthly payment on 240,000 over 25 years at 5.4 percent, and how much interest is that in total?" With the agent above, the session records this sequence:
event 1 user : the question
event 2 agent : ```python
P, r, n = 240000, 0.054/12, 25*12
pay = P*r/(1-(1+r)**-n)
print(round(pay, 2), round(pay*n - P, 2))
```
event 3 model : codeExecutionResult stdout="1459.51 197853.54"
stateDelta: _code_execution_context ..., error count reset
event 4 agent : "The payment is about 1,459.51 a month; total interest is
about 197,854 over the 300 payments."If the model had written pay * n - P as a final bare expression without print, stdout would be empty and the model would see a blank result. Most models then rewrite with a print, which costs one model call. A user who sees the second answer never knows. Your audit log does, and a high rate of empty-stdout runs is a prompt bug worth fixing.
Failure modes
- Shared-mode container in production. Generated code can reach the network and the cloud metadata endpoint. Turn strict mode on and verify it in a startup check.
- Warnings spend the retry budget. Two noisy runs disable execution for the invocation. Silence warnings and classify stderr.
- Expecting files from the container executor. Charts and CSVs are not returned in 1.11.0. Use Vertex or a custom executor that collects files.
- Vertex executor without a resource name. It logs that the interpreter is unavailable and returns empty output, so the model sees silence rather than an error. Fail fast at startup if the resource name is missing.
- State assumptions. The model defines a variable in one block and uses it in the next; the second run fails with
NameError. Say "each block runs alone" in the instruction. - Huge stdout. Every byte goes back into the prompt. Cap it in a wrapper.
- Leaked containers. Forgetting
close()leaves containers behind on redeploys. Register the executor with your shutdown hooks.
Trade-offs
The built-in executor is the least work and the least control: no image to maintain, no audit hook, Gemini only. The container executor gives you control and an audit trail at the cost of operating Docker, keeping strict mode on and handling cold-start latency, since strict mode creates a container per run. The Vertex executor returns files and runs off your hosts, but depends on a beta API. Across all of them, every execution is at least one extra model call, so a code executor raises latency and cost per answer in exchange for exactness. Use it on agents whose answers are numbers, and leave it off agents whose answers are prose. Model calls and tool dispatch are explained in how ADK Java dispatches tools; where the saved files end up is covered in ADK Java artifacts.
What to do next
- Decide whether the agent needs computation at all. If its answers are not numbers or tables, do not attach an executor.
- Pick the executor: built-in for Gemini-only prototypes, strict container for control, Vertex if you need files back.
- For containers, build a pinned image with only the libraries you promise in the instruction, and call
setStrictSandbox(true). - Write the instruction rules: print every answer, each block runs alone, never state an unprinted number.
- Wrap the executor to cap output, classify stderr and write an audit record of code, output and timing.
- Configure a durable artifact service if any executor returns files.
- Add alerts on execution error rate, empty-stdout rate and execution time, and close the executor on shutdown.
- Run a prompt-injection test from the sandboxing guide against your configuration before release.