A language model remembers nothing between calls. What feels like an agent's short-term memory is the request your framework rebuilds before every model call: the instructions, the conversation so far, the tool schemas and the tool results. In ADK for Java that request is assembled from the session's event log, the session's state map and the agent definition. Whatever is in it, the model knows; whatever is not, the model has forgotten, however recently it happened.
This article shows exactly how ADK Java builds that request, how to measure what it costs per turn, and the levers that keep it small: trimming tool output at the source, moving facts into state, running sub-agents without history, and compaction. API names were checked against the google/adk-java main branch on 2026-10-04; check the Javadoc of the release you run before copying a builder call. Compaction itself is covered in ADK Java context compression and the state scopes in session context in ADK Java.
Three routes into the context window
There is no class called ShortTermMemory in ADK Java. Within one session, information reaches the model by one of three routes, and each has a different cost profile.
| Route | What lives there | How it reaches the model | Cost grows with |
|---|---|---|---|
| Session events | user messages, model replies, function calls and responses | converted to request contents every call | every turn, forever, unless compacted |
| Session state | small facts: ids, preferences, extracted slots | only where an instruction says {key} | nothing, unless you inject it |
| Agent definition | instruction, global instruction, tool declarations | every call, unchanged | number and size of tools |
Long-term memory, the MemoryService that searches past sessions, is a fourth route that only fires when a tool or callback asks it to; see the MemoryService article. This page is about the first three, because they decide the token bill of every single turn.
How a request is assembled
Request building runs as a chain of processors. The one that matters most for memory is the contents processor, which walks the session's events and turns them into the list of Content objects the model sees. On main it does five things worth knowing.
- It skips empty events: no content, no parts, or parts that are invisible, such as thoughts without a function call.
- It filters by branch. A parallel agent gives each child its own branch, so parallel siblings do not see each other's work; agents run in sequence share a branch and do see earlier turns.
- It rewrites other agents' turns as attributed, quoted context rather than replaying them as if this agent had said them. Their text, tool calls and tool results are labelled with the agent name.
- It substitutes compaction summaries: events whose timestamps fall inside a compaction event's range are left out and the summary goes in their place.
- If the agent is built with
includeContents(LlmAgent.IncludeContents.NONE), it returns only the current turn: it scans back to the latest user message or hand-off and starts there.
Separately, the instruction is resolved against state. {key} is replaced by the state value, {user:tier} reads a user-scoped key, {key?} becomes an empty string when missing, and {artifact.name} loads an artifact. A missing non-optional key fails the call with a Context variable not found error, which is the most common first-day surprise.
Budgeting the window, worked through
Treat the context window as a budget with four lines. Fixed: instructions and tool declarations, paid on every call. History: everything the contents processor keeps. Current turn: the new user message and, inside a tool loop, each call and result so far. Reserve: room for the answer and for one more tool result.
Work through an order-support agent with illustrative numbers. Its instruction is 600 tokens and its six tools declare 1,400 tokens of schema, so 2,000 tokens are spent before the user types. A typical turn adds a 60-token question, a function call of 40 tokens, a function response and a 150-token reply. If the tool returns the full order record as JSON, the response is about 3,000 tokens and the turn costs about 3,250. After twenty turns the history is 65,000 tokens and every new call pays for all of it, even though the model needed only the status and delivery date from each record. Trim the tool output to four fields, about 80 tokens, and the same twenty turns cost about 6,600 tokens of history. Nothing else changed.
Two lessons follow. Tool responses, not chat, dominate history in most agents. And a model call inside a tool loop pays for the history once per step, so an invocation that calls three tools in sequence reads the history four times.
Keep payloads out of the log
The cheapest token is one that never enters the log. Return what the model needs to decide the next step, and keep everything else in your own systems keyed by an id that the model can pass back. FunctionTool.create wraps a static method; a parameter named toolContext receives the ToolContext, which is how the tool writes small facts into state.
public class OrderTools {
public static Map<String, Object> fetchOrder(
@Schema(name = "orderId", description = "Order id such as ORD-1042") String orderId,
ToolContext toolContext) {
Order order = orderStore.get(orderId); // your own data access
toolContext.state().put("last_order_id", orderId); // a fact, not a payload
return Map.of(
"status", order.status(),
"carrier", order.carrier(),
"eta", order.eta().toString(),
"lineCount", order.lines().size()); // not the 200 line items
}
}
LlmAgent support = LlmAgent.builder()
.name("order_support")
.model(MODEL_ID) // the model you deploy
.instruction("You help customers with their orders. Customer tier: {user:tier?}. "
+ "Order under discussion: {last_order_id?}. Answer in plain sentences.")
.tools(FunctionTool.create(OrderTools.class, "fetchOrder"))
.build();The state write matters as much as the trimming. Once last_order_id is in state and injected into the instruction, the model knows which order is in play even after compaction has summarised away the turn where it was mentioned. The rule of thumb is: identifiers and decisions go in state, evidence stays in your systems, and the event log carries only what the model must reason over. Note the ? on both placeholders; the first turn has neither key yet, and without it the call would fail.
Measuring what the model actually saw
You cannot manage a budget you do not read. Each model response event carries usage metadata, and promptTokenCount() is the size of the request the model actually received, after every processor ran. Record it per agent and per invocation as events stream out of the runner.
runner.runAsync(userId, sessionId, Content.fromParts(Part.fromText(userText)))
.doOnNext(event -> event.usageMetadata()
.flatMap(usage -> usage.promptTokenCount())
.ifPresent(tokens -> metrics.record(
event.author(), event.invocationId(), tokens)))
.blockingSubscribe(event -> handle(event));Then watch three numbers. Prompt tokens per call, against the window of the model you call: alert well before you reach it, for example at 60 to 70 percent, because one large tool result can add tens of thousands of tokens in a single step. Growth per invocation: the slope tells you how many turns a session can run before compaction must fire. Calls per invocation: a tool loop that takes five steps multiplies the history cost by five. To audit an old session, load it from the session service and read the same field from each stored event, provided your session service persists usage metadata; the in-memory service keeps the event objects as they were, but check a custom store.
Sub-agents without history
Not every agent needs the conversation. A classifier that labels the latest message, an extractor that pulls an order id, or a formatter that rewrites the final answer can run on the current turn alone. Build it with IncludeContents.NONE and have it hand its result to the next agent through state with outputKey, which stores the agent's final text under that key.
LlmAgent intent = LlmAgent.builder()
.name("intent_classifier")
.model(MODEL_ID)
.includeContents(LlmAgent.IncludeContents.NONE)
.instruction("Label the latest user message as status, return or other. "
+ "Reply with the label only.")
.outputKey("intent")
.build();
SequentialAgent pipeline = SequentialAgent.builder()
.name("support_pipeline")
.subAgents(intent, support) // support reads {intent?} in its instruction
.build();The classifier now costs a few hundred tokens whatever the session length. The trade-off is real: a message like yes, that one is meaningless without history, so use NONE only for steps whose input is self-contained, and give the step the state it needs through placeholders instead of the transcript.
The same thinking applies to agent trees with several specialists. Every agent that runs with default contents pays for the history it can see, and every hand-off adds events that later agents read as quoted context. A tree of four specialists that each see the full transcript can cost four times what a single agent would. Before adding an agent, ask what it needs to read. Often the answer is the current message and two state keys, and an agent built that way stays cheap no matter how long the session runs. Write that contract down in the agent's instruction so the next person to edit it does not quietly widen it.
What callbacks are for here
A before-model callback sees the request last, so it is tempting to make it the place where history is trimmed. On main, BeforeModelCallbackSync receives a CallbackContext and an LlmRequest.Builder and returns Optional<LlmResponse>: empty to call the model, or a response to skip the call entirely. The builder has setters, including contents(...) and appendInstructions(...), but no getter for the contents it already holds, so rewriting history there is awkward and easy to get wrong. Use it for what it does well.
.beforeModelCallbackSync((callbackContext, request) -> {
Object plan = callbackContext.state().get("user:plan");
if ("free".equals(plan)) {
request.appendInstructions(List.of("Keep answers under 80 words."));
}
return Optional.empty(); // empty means: call the model as usual
})Trim at the source, in tools, and bound history with compaction. Callbacks are for guards, small instruction changes and short-circuiting; more patterns are in the ADK Java callbacks article.
Failure modes
What goes wrong, and what it looks like:
- The agent forgets a fact from ten turns ago. It was only in the log and compaction summarised it away, or a NONE agent never saw it. Put the fact in state and inject it.
- Context variable not found on the first turn. A placeholder without
?names a key that only a later step writes. - Cost climbs linearly per session. Large tool responses are accumulating. Look at the biggest function response events first.
- The window overflows mid-loop. One tool returned a document. Cap tool output size in the tool itself and return a handle instead.
- A sub-agent repeats work a sibling did. The sibling ran in parallel on its own branch, or the agent was built with NONE. Pass results through state keys, not through the transcript.
- temp: values are still visible next turn, or vanish when you expected them. Their persistence differs by session service; never rely on either outcome.
Trade-offs between the levers
| Lever | Saves | Costs |
|---|---|---|
| Trim tool output at the source | the largest share of history | a second call when the model needs detail |
| Facts into state, injected by {key} | re-reading old turns to find ids | you must choose which facts matter |
| IncludeContents.NONE sub-agents | all history for that step | no access to earlier turns |
| Compaction | bounded history on long sessions | summary calls and lossy summaries |
| Fewer, smaller tool schemas | fixed cost on every call | possibly more agents to route between |
Apply them roughly in that order. The first two are free and prevent most problems; compaction is the safety net, not the design.
What to do next
- Record promptTokenCount per agent and invocation for one day of real traffic.
- Find the five largest function response events and cut each tool to the fields the model uses.
- Move ids and decisions into state and reference them with optional placeholders in instructions.
- Mark self-contained steps with IncludeContents.NONE and connect them with outputKey.
- Configure compaction with a threshold at 60 to 70 percent of your model window, then confirm the slope flattens.
- Add an alert on prompt tokens per call and on calls per invocation.