When an ADK Java agent streams, your client does not receive text. It receives a sequence of Event objects on an RxJava Flowable. Some are partial, some are aggregates of what was just streamed, and some carry tool calls, tool results, errors or control flags. Rendering them correctly, persisting the right ones and turning them into a wire protocol a browser can consume is where most streaming bugs live. The classic examples are text that appears twice, tool calls that a client tries to act on too early, and a reconnecting user who sees a different answer from the one that streamed.
This page is about the events themselves: which ones each StreamingMode produces, which fields to read while streaming, what the runner persists to the session, and how to translate the stream into Server-Sent Events with a client that cannot double-render. API names and behaviour were read from the google/adk-java sources on main on 2026-10-03, and method names were confirmed against the compiled 1.10.1 artifact. Treat anything not named here as unverified. The broader streaming architecture (backpressure, threads, cancellation) is covered in the streaming architecture page, and the general Event contract in ADK Java events.
Three modes, three event shapes
RunConfig.StreamingMode has three values: NONE (the default), SSE and BIDI. In the request path, BaseLlmFlow calls the model with llm.generateContent(request, runConfig.streamingMode() == StreamingMode.SSE). So only SSE turns on token streaming for runAsync. Live, bidirectional sessions go through a different entry point, Runner.runLive, which takes a LiveRequestQueue instead of a single message.
| Entry point and mode | What one model call produces | Use it for |
|---|---|---|
runAsync, NONE | One non-partial event per model response | Batch jobs, back-end callers, tests, anything that waits for the whole answer |
runAsync, SSE | Partial events as chunks arrive, then one non-partial aggregate | Chat UIs and any client that should show text as it is generated |
runLive (typically with BIDI) | Events from a live model connection while the client keeps sending | Voice and real-time sessions where input and output overlap |
Whatever the mode, the subscriber sees one Flowable<Event>, and tool calls, tool results and agent transfers arrive on it as events too. What changes is how many events carry text and whether they are marked partial. Switching a client from NONE to SSE without changing how it handles events is the most common way to introduce the double-render bug.
The fields that matter while streaming
An Event has many accessors. While streaming, a handful decide what a consumer should do. All the names below exist on com.google.adk.events.Event.
| Accessor | Type | What it tells a streaming consumer |
|---|---|---|
partial() | Optional<Boolean> | True on in-progress chunks. Render, never persist, never act on |
finalResponse() | boolean | True when the event is a settled answer: no function calls or responses, not partial (or a long-running tool is pending, or summarisation is skipped) |
functionCalls() | ImmutableList<FunctionCall> | The tools the model wants run. Display "calling X"; the framework executes them |
functionResponses() | ImmutableList<FunctionResponse> | Tool results. Show progress or a summary, never raw payloads to end users by default |
errorCode() / errorMessage() | Optional<FinishReason> / Optional<String> | A model-level error that arrived as an event, not as a stream error |
turnComplete() / interrupted() | Optional<Boolean> | Live-session flags; check the ADK documentation for their exact meaning |
author(), invocationId(), id() | String | Which agent spoke, which run it belongs to, and the event's identity |
The body of finalResponse() is worth knowing exactly. It returns true if skipSummarization is set or a long-running tool call is pending. Otherwise it returns true only when the event has no function calls, no function responses, is not partial, and has no trailing code-execution result. That is why a client should key its "answer is done" logic on finalResponse() rather than on stream completion. A multi-agent run can produce several final responses, from different authors, before the Flowable completes.
Inside SSE mode: chunk, partial, aggregate, persist
In SSE mode, the Gemini model adapter feeds the streamed chunks through a StreamingResponseAggregator. Each non-empty chunk becomes an LlmResponse marked partial(true). When the model stream ends, the aggregator emits one more response with the accumulated parts and no partial flag. The source notes that this final response is emitted even without a finish reason, so accumulated content is never dropped.
The flow then treats the two kinds differently. In BaseLlmFlow, a model event is passed straight through, without running tools, if it has no function calls or is partial. Only a non-partial event with function calls triggers runFunctionCalls. So a client that sees a function call in a partial event and tries to act on it is ahead of the framework, and the call will be executed once, later, from the aggregate. The run loop ends when the last event of a step is a final response, or carries an endInvocation action.
The runner applies the last filter. Partial events are emitted to the caller but not passed to sessionService.appendEvent; non-partial events are appended. The consequence is simple and important: the session history contains the aggregate, never the chunks. Anything you rebuild from the session (a page reload, a resumed conversation, an audit) sees the settled text, which may differ from the concatenated deltas in whitespace or chunk boundaries. Never treat client-side accumulated deltas as the record.
A gateway: Event to Server-Sent Events
Browsers consume Server-Sent Events natively. The ADK dev server already exposes POST /run_sse with a body carrying appName, userId, sessionId, newMessage, streaming and stateDelta; it is a good reference but is a development tool. For production you usually want your own endpoint that authenticates the user, owns the session ids and emits a small, typed protocol instead of raw event JSON. Here is one with Spring MVC's SseEmitter:
@RestController
class ChatStreamController {
private final Runner runner; // e.g. new InMemoryRunner(rootAgent) in development
ChatStreamController(Runner runner) { this.runner = runner; }
@PostMapping(path = "/chat/{sessionId}", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
SseEmitter chat(@PathVariable String sessionId, @RequestBody String text, Principal user) {
SseEmitter emitter = new SseEmitter(120_000L);
RunConfig cfg = RunConfig.builder()
.setStreamingMode(RunConfig.StreamingMode.SSE)
.setMaxLlmCalls(20) // the session must exist, or use setAutoCreateSession(true)
.build();
AtomicLong seq = new AtomicLong();
Disposable sub = runner
.runAsync(user.getName(), sessionId, Content.fromParts(Part.fromText(text)), cfg)
.subscribe(
ev -> send(emitter, seq.incrementAndGet(), ev),
err -> { frame(emitter, seq.incrementAndGet(), "error", "{\"retryable\":true}");
emitter.complete(); },
() -> { frame(emitter, seq.incrementAndGet(), "done", "{}");
emitter.complete(); });
emitter.onCompletion(sub::dispose);
emitter.onTimeout(sub::dispose);
return emitter;
}
private void send(SseEmitter em, long n, Event ev) {
String text = ev.content().flatMap(Content::parts).orElse(List.of()).stream()
.map(part -> part.text().orElse(""))
.collect(Collectors.joining());
if (ev.errorCode().isPresent()) {
frame(em, n, "error", Json.of("code", ev.errorCode().get().toString()));
} else if (ev.partial().orElse(false)) {
if (!text.isEmpty()) frame(em, n, "delta", Json.of("author", ev.author(), "text", text));
} else if (!ev.functionCalls().isEmpty()) {
if (!text.isEmpty()) frame(em, n, "message", Json.of("author", ev.author(), "text", text, "final", false));
frame(em, n, "tool_call", Json.of("names", ev.functionCalls().stream()
.map(fc -> fc.name().orElse("?")).toList()));
} else if (!ev.functionResponses().isEmpty()) {
frame(em, n, "tool_result", Json.of("count", ev.functionResponses().size()));
} else if (!text.isEmpty()) {
frame(em, n, "message", Json.of("author", ev.author(), "id", ev.id(),
"text", text, "final", ev.finalResponse()));
}
}
private void frame(SseEmitter em, long n, String type, String json) {
try {
em.send(SseEmitter.event().id(Long.toString(n)).name(type).data(json));
} catch (IOException e) {
em.completeWithError(e); // client went away; onCompletion disposes the run
}
}
}Json.of stands in for your JSON helper, and the error payload is deliberately generic. Four design choices carry the weight here. Frames are typed by event kind, so the browser never has to inspect model internals. Tool results are summarised, not forwarded, because raw tool payloads can contain data the user should not see. Disposal is wired to completion and timeout, so a closed tab cancels the upstream Flowable instead of leaving a model call running. And the sequence id is the gateway's own counter, because nothing in the event contract promises ids suited to SSE Last-Event-ID resumption. If you run on Spring, ADK Java with Spring covers wiring the runner as a bean.
The client: append, then replace
The client rule is: append deltas to a draft, then replace the draft with the message when it arrives. Never append both. A reducer makes that impossible to get wrong:
// state: { messages: Message[], draft: { author: string, text: string } | null }
function reduce(state, frame) {
const d = JSON.parse(frame.data);
switch (frame.type) {
case "delta":
if (!state.draft || state.draft.author !== d.author) {
return { ...state, draft: { author: d.author, text: d.text } };
}
return { ...state, draft: { ...state.draft, text: state.draft.text + d.text } };
case "message": // the aggregate replaces the draft; it is the record
return { ...state, messages: [...state.messages, d], draft: null };
case "tool_call":
return { ...state, draft: null, status: "Calling " + d.names.join(", ") };
case "error":
case "done":
return { ...state, draft: null, status: frame.type };
default:
return state;
}
}Keying the draft by author matters in multi-agent runs, where a sub-agent's partials can follow the coordinator's aggregate. Clearing the draft on errors and on completion matters for the other common bug: a half-rendered sentence that stays on screen after the run failed. On reconnect, discard drafts and reload messages from your own history endpoint, which reads the session. Do not try to replay deltas.
Live sessions with LiveRequestQueue
Live sessions invert the shape. Instead of one message in and a stream out, the client keeps pushing input through a LiveRequestQueue while events stream back. The queue's public methods are content(Content) for text turns, realtime(Blob) for media such as audio chunks, send(LiveRequest), get() and close(). Internally it is a serialised RxJava MulticastProcessor, so it is safe to push from a WebSocket handler thread.
LiveRequestQueue queue = new LiveRequestQueue();
RunConfig liveCfg = RunConfig.builder()
.setStreamingMode(RunConfig.StreamingMode.BIDI)
.build();
Disposable sub = runner.runLive(userId, sessionId, queue, liveCfg)
.subscribe(ev -> socket.sendEvent(ev), socket::fail, socket::close);
socket.onText(t -> queue.content(Content.fromParts(Part.fromText(t))));
socket.onAudio(blob -> queue.realtime(blob));
socket.onClose(() -> { queue.close(); sub.dispose(); });Live sessions need a model and model adapter that support live connections, and their events can carry the turnComplete() and interrupted() flags; check the ADK documentation for exactly when each is set before driving UI state from them. Confirm the exact live-capable model names and audio formats against the current ADK and Gemini documentation before you build on them; they change faster than the event contract.
Worked example: one tool call, streamed
A user asks a support agent with one tool, lookupOrder, "Where is order 1042?" in SSE mode. The gateway emits the following frames:
id:1 event:tool_call {"names":["lookupOrder"]}
id:2 event:tool_result {"count":1}
id:3 event:delta {"author":"support","text":"Order 1042 shipped "}
id:4 event:delta {"author":"support","text":"on 30 September and "}
id:5 event:delta {"author":"support","text":"is due Monday."}
id:6 event:message {"author":"support","id":"...","text":"Order 1042 shipped on 30 September and is due Monday.","final":true}
id:7 event:done {}The model's first response was a function call. Depending on how the model streams, that call may also have appeared in a partial chunk; the gateway shows it only from the aggregate, and the flow runs the tool exactly once from the aggregate too. The session now holds the user message, the function-call event, the function-response event and the final text event, and none of the three deltas. If the user reloads after frame 5, the client discards the draft, fetches history and sees the user message and the tool call but no answer text until the aggregate is persisted. That is correct: frame 6 is what was said.
Failure modes
- Double rendering. The client appends deltas and then appends the aggregate. Fix: replace on the non-partial message, as in the reducer.
- Acting on partial function calls. A UI or proxy triggers side effects from a partial event. The framework only executes calls from the non-partial event, so any such action is a duplicate.
- Treating stream completion as the answer. Multi-agent runs produce several events with
finalResponse()true; decide per author and per turn. - Errors that are events. A model error can arrive as an event with
errorCode()set while the Flowable completes normally. Check the field; do not rely only ononError. - Runaway loops. A tool that always triggers another call keeps the stream open.
maxLlmCallsdefaults to 500; set a budget suited to an interactive request. - Orphaned runs. A closed tab with no disposal wiring leaves model calls and tools running. Dispose the subscription on completion, timeout and send failure.
- Leaking tool payloads. Forwarding raw events puts function arguments and responses in the browser. Summarise them in the gateway, and see ADK Java callbacks for filtering at the agent level.
Trade-offs
SSE mode improves time to first token and perceived latency, at the cost of more events, more network frames and a client that must handle two kinds of text event. NONE mode is simpler and fits back-end callers, and its single aggregate per model call is exactly what SSE mode persists, so session history is the same either way. Live mode is the only option for overlapping input and output, and it brings its own model and transport constraints. A typed gateway protocol costs a little code but decouples browsers from ADK's event schema, so upgrading the framework does not break deployed clients.
What to do next
- List every client that consumes your agent's stream and check that each replaces the draft on the non-partial event rather than appending it.
- Audit side-effecting code paths for any that read function calls from partial events, and move them to the aggregate or to the tool itself.
- Put a gateway in front of runAsync that emits typed frames, summarises tool results, disposes on completion and timeout, and numbers frames itself.
- Set maxLlmCalls for interactive requests and alert on runs that hit it.
- Add a test that streams a canned model response and asserts the frame sequence, and that the session contains the aggregate but no partials.
- If you need voice or interruptible sessions, prototype runLive with a LiveRequestQueue after confirming live-capable models in the current documentation.