When an ADK Java agent streams, your client does not receive text. It receives a sequence of Event objects on an RxJava Flowable. Some are partial, some are aggregates of what was just streamed, and some carry tool calls, tool results, errors or control flags. Rendering them correctly, persisting the right ones and turning them into a wire protocol a browser can consume is where most streaming bugs live. The classic examples are text that appears twice, tool calls that a client tries to act on too early, and a reconnecting user who sees a different answer from the one that streamed.

This page is about the events themselves: which ones each StreamingMode produces, which fields to read while streaming, what the runner persists to the session, and how to translate the stream into Server-Sent Events with a client that cannot double-render. API names and behaviour were read from the google/adk-java sources on main on 2026-10-03, and method names were confirmed against the compiled 1.10.1 artifact. Treat anything not named here as unverified. The broader streaming architecture (backpressure, threads, cancellation) is covered in the streaming architecture page, and the general Event contract in ADK Java events.

Three modes, three event shapes

RunConfig.StreamingMode has three values: NONE (the default), SSE and BIDI. In the request path, BaseLlmFlow calls the model with llm.generateContent(request, runConfig.streamingMode() == StreamingMode.SSE). So only SSE turns on token streaming for runAsync. Live, bidirectional sessions go through a different entry point, Runner.runLive, which takes a LiveRequestQueue instead of a single message.

Entry point and modeWhat one model call producesUse it for
runAsync, NONEOne non-partial event per model responseBatch jobs, back-end callers, tests, anything that waits for the whole answer
runAsync, SSEPartial events as chunks arrive, then one non-partial aggregateChat UIs and any client that should show text as it is generated
runLive (typically with BIDI)Events from a live model connection while the client keeps sendingVoice and real-time sessions where input and output overlap

Whatever the mode, the subscriber sees one Flowable<Event>, and tool calls, tool results and agent transfers arrive on it as events too. What changes is how many events carry text and whether they are marked partial. Switching a client from NONE to SSE without changing how it handles events is the most common way to introduce the double-render bug.

Model APIchunksGemini.generateContentstream = (mode == SSE)BaseLlmFlowtools on non-partialRunnerpersist non-partialSessionServiceappendEventfinal onlyYour gatewayEvent -> SSE framesevery eventBrowser reducerappend delta, replace on finaltext/event-streamSSE mode, one model call:partial(true) chunkpartial(true) chunk...aggregated, partial unset -> persistedNONE mode: only the aggregated eventrunLive: events from a live connectionPartials reach the caller but never the session; function calls run only from the non-partial event.
Where streaming events come from and where they go in ADK Java SSE mode.

The fields that matter while streaming

An Event has many accessors. While streaming, a handful decide what a consumer should do. All the names below exist on com.google.adk.events.Event.

AccessorTypeWhat it tells a streaming consumer
partial()Optional<Boolean>True on in-progress chunks. Render, never persist, never act on
finalResponse()booleanTrue when the event is a settled answer: no function calls or responses, not partial (or a long-running tool is pending, or summarisation is skipped)
functionCalls()ImmutableList<FunctionCall>The tools the model wants run. Display "calling X"; the framework executes them
functionResponses()ImmutableList<FunctionResponse>Tool results. Show progress or a summary, never raw payloads to end users by default
errorCode() / errorMessage()Optional<FinishReason> / Optional<String>A model-level error that arrived as an event, not as a stream error
turnComplete() / interrupted()Optional<Boolean>Live-session flags; check the ADK documentation for their exact meaning
author(), invocationId(), id()StringWhich agent spoke, which run it belongs to, and the event's identity

The body of finalResponse() is worth knowing exactly. It returns true if skipSummarization is set or a long-running tool call is pending. Otherwise it returns true only when the event has no function calls, no function responses, is not partial, and has no trailing code-execution result. That is why a client should key its "answer is done" logic on finalResponse() rather than on stream completion. A multi-agent run can produce several final responses, from different authors, before the Flowable completes.

Inside SSE mode: chunk, partial, aggregate, persist

In SSE mode, the Gemini model adapter feeds the streamed chunks through a StreamingResponseAggregator. Each non-empty chunk becomes an LlmResponse marked partial(true). When the model stream ends, the aggregator emits one more response with the accumulated parts and no partial flag. The source notes that this final response is emitted even without a finish reason, so accumulated content is never dropped.

The flow then treats the two kinds differently. In BaseLlmFlow, a model event is passed straight through, without running tools, if it has no function calls or is partial. Only a non-partial event with function calls triggers runFunctionCalls. So a client that sees a function call in a partial event and tries to act on it is ahead of the framework, and the call will be executed once, later, from the aggregate. The run loop ends when the last event of a step is a final response, or carries an endInvocation action.

The runner applies the last filter. Partial events are emitted to the caller but not passed to sessionService.appendEvent; non-partial events are appended. The consequence is simple and important: the session history contains the aggregate, never the chunks. Anything you rebuild from the session (a page reload, a resumed conversation, an audit) sees the settled text, which may differ from the concatenated deltas in whitespace or chunk boundaries. Never treat client-side accumulated deltas as the record.

A gateway: Event to Server-Sent Events

Browsers consume Server-Sent Events natively. The ADK dev server already exposes POST /run_sse with a body carrying appName, userId, sessionId, newMessage, streaming and stateDelta; it is a good reference but is a development tool. For production you usually want your own endpoint that authenticates the user, owns the session ids and emits a small, typed protocol instead of raw event JSON. Here is one with Spring MVC's SseEmitter:

@RestController
class ChatStreamController {
  private final Runner runner;   // e.g. new InMemoryRunner(rootAgent) in development

  ChatStreamController(Runner runner) { this.runner = runner; }

  @PostMapping(path = "/chat/{sessionId}", produces = MediaType.TEXT_EVENT_STREAM_VALUE)
  SseEmitter chat(@PathVariable String sessionId, @RequestBody String text, Principal user) {
    SseEmitter emitter = new SseEmitter(120_000L);
    RunConfig cfg = RunConfig.builder()
        .setStreamingMode(RunConfig.StreamingMode.SSE)
        .setMaxLlmCalls(20)          // the session must exist, or use setAutoCreateSession(true)
        .build();
    AtomicLong seq = new AtomicLong();

    Disposable sub = runner
        .runAsync(user.getName(), sessionId, Content.fromParts(Part.fromText(text)), cfg)
        .subscribe(
            ev -> send(emitter, seq.incrementAndGet(), ev),
            err -> { frame(emitter, seq.incrementAndGet(), "error", "{\"retryable\":true}");
                     emitter.complete(); },
            () -> { frame(emitter, seq.incrementAndGet(), "done", "{}");
                    emitter.complete(); });

    emitter.onCompletion(sub::dispose);
    emitter.onTimeout(sub::dispose);
    return emitter;
  }

  private void send(SseEmitter em, long n, Event ev) {
    String text = ev.content().flatMap(Content::parts).orElse(List.of()).stream()
        .map(part -> part.text().orElse(""))
        .collect(Collectors.joining());
    if (ev.errorCode().isPresent()) {
      frame(em, n, "error", Json.of("code", ev.errorCode().get().toString()));
    } else if (ev.partial().orElse(false)) {
      if (!text.isEmpty()) frame(em, n, "delta", Json.of("author", ev.author(), "text", text));
    } else if (!ev.functionCalls().isEmpty()) {
      if (!text.isEmpty()) frame(em, n, "message", Json.of("author", ev.author(), "text", text, "final", false));
      frame(em, n, "tool_call", Json.of("names", ev.functionCalls().stream()
          .map(fc -> fc.name().orElse("?")).toList()));
    } else if (!ev.functionResponses().isEmpty()) {
      frame(em, n, "tool_result", Json.of("count", ev.functionResponses().size()));
    } else if (!text.isEmpty()) {
      frame(em, n, "message", Json.of("author", ev.author(), "id", ev.id(),
          "text", text, "final", ev.finalResponse()));
    }
  }

  private void frame(SseEmitter em, long n, String type, String json) {
    try {
      em.send(SseEmitter.event().id(Long.toString(n)).name(type).data(json));
    } catch (IOException e) {
      em.completeWithError(e);   // client went away; onCompletion disposes the run
    }
  }
}

Json.of stands in for your JSON helper, and the error payload is deliberately generic. Four design choices carry the weight here. Frames are typed by event kind, so the browser never has to inspect model internals. Tool results are summarised, not forwarded, because raw tool payloads can contain data the user should not see. Disposal is wired to completion and timeout, so a closed tab cancels the upstream Flowable instead of leaving a model call running. And the sequence id is the gateway's own counter, because nothing in the event contract promises ids suited to SSE Last-Event-ID resumption. If you run on Spring, ADK Java with Spring covers wiring the runner as a bean.

The client: append, then replace

The client rule is: append deltas to a draft, then replace the draft with the message when it arrives. Never append both. A reducer makes that impossible to get wrong:

// state: { messages: Message[], draft: { author: string, text: string } | null }
function reduce(state, frame) {
  const d = JSON.parse(frame.data);
  switch (frame.type) {
    case "delta":
      if (!state.draft || state.draft.author !== d.author) {
        return { ...state, draft: { author: d.author, text: d.text } };
      }
      return { ...state, draft: { ...state.draft, text: state.draft.text + d.text } };
    case "message":   // the aggregate replaces the draft; it is the record
      return { ...state, messages: [...state.messages, d], draft: null };
    case "tool_call":
      return { ...state, draft: null, status: "Calling " + d.names.join(", ") };
    case "error":
    case "done":
      return { ...state, draft: null, status: frame.type };
    default:
      return state;
  }
}

Keying the draft by author matters in multi-agent runs, where a sub-agent's partials can follow the coordinator's aggregate. Clearing the draft on errors and on completion matters for the other common bug: a half-rendered sentence that stays on screen after the run failed. On reconnect, discard drafts and reload messages from your own history endpoint, which reads the session. Do not try to replay deltas.

Live sessions with LiveRequestQueue

Live sessions invert the shape. Instead of one message in and a stream out, the client keeps pushing input through a LiveRequestQueue while events stream back. The queue's public methods are content(Content) for text turns, realtime(Blob) for media such as audio chunks, send(LiveRequest), get() and close(). Internally it is a serialised RxJava MulticastProcessor, so it is safe to push from a WebSocket handler thread.

LiveRequestQueue queue = new LiveRequestQueue();
RunConfig liveCfg = RunConfig.builder()
    .setStreamingMode(RunConfig.StreamingMode.BIDI)
    .build();

Disposable sub = runner.runLive(userId, sessionId, queue, liveCfg)
    .subscribe(ev -> socket.sendEvent(ev), socket::fail, socket::close);

socket.onText(t -> queue.content(Content.fromParts(Part.fromText(t))));
socket.onAudio(blob -> queue.realtime(blob));
socket.onClose(() -> { queue.close(); sub.dispose(); });

Live sessions need a model and model adapter that support live connections, and their events can carry the turnComplete() and interrupted() flags; check the ADK documentation for exactly when each is set before driving UI state from them. Confirm the exact live-capable model names and audio formats against the current ADK and Gemini documentation before you build on them; they change faster than the event contract.

Worked example: one tool call, streamed

A user asks a support agent with one tool, lookupOrder, "Where is order 1042?" in SSE mode. The gateway emits the following frames:

id:1  event:tool_call   {"names":["lookupOrder"]}
id:2  event:tool_result {"count":1}
id:3  event:delta       {"author":"support","text":"Order 1042 shipped "}
id:4  event:delta       {"author":"support","text":"on 30 September and "}
id:5  event:delta       {"author":"support","text":"is due Monday."}
id:6  event:message     {"author":"support","id":"...","text":"Order 1042 shipped on 30 September and is due Monday.","final":true}
id:7  event:done        {}

The model's first response was a function call. Depending on how the model streams, that call may also have appeared in a partial chunk; the gateway shows it only from the aggregate, and the flow runs the tool exactly once from the aggregate too. The session now holds the user message, the function-call event, the function-response event and the final text event, and none of the three deltas. If the user reloads after frame 5, the client discards the draft, fetches history and sees the user message and the tool call but no answer text until the aggregate is persisted. That is correct: frame 6 is what was said.

Failure modes

  • Double rendering. The client appends deltas and then appends the aggregate. Fix: replace on the non-partial message, as in the reducer.
  • Acting on partial function calls. A UI or proxy triggers side effects from a partial event. The framework only executes calls from the non-partial event, so any such action is a duplicate.
  • Treating stream completion as the answer. Multi-agent runs produce several events with finalResponse() true; decide per author and per turn.
  • Errors that are events. A model error can arrive as an event with errorCode() set while the Flowable completes normally. Check the field; do not rely only on onError.
  • Runaway loops. A tool that always triggers another call keeps the stream open. maxLlmCalls defaults to 500; set a budget suited to an interactive request.
  • Orphaned runs. A closed tab with no disposal wiring leaves model calls and tools running. Dispose the subscription on completion, timeout and send failure.
  • Leaking tool payloads. Forwarding raw events puts function arguments and responses in the browser. Summarise them in the gateway, and see ADK Java callbacks for filtering at the agent level.

Trade-offs

SSE mode improves time to first token and perceived latency, at the cost of more events, more network frames and a client that must handle two kinds of text event. NONE mode is simpler and fits back-end callers, and its single aggregate per model call is exactly what SSE mode persists, so session history is the same either way. Live mode is the only option for overlapping input and output, and it brings its own model and transport constraints. A typed gateway protocol costs a little code but decouples browsers from ADK's event schema, so upgrading the framework does not break deployed clients.

What to do next

  1. List every client that consumes your agent's stream and check that each replaces the draft on the non-partial event rather than appending it.
  2. Audit side-effecting code paths for any that read function calls from partial events, and move them to the aggregate or to the tool itself.
  3. Put a gateway in front of runAsync that emits typed frames, summarises tool results, disposes on completion and timeout, and numbers frames itself.
  4. Set maxLlmCalls for interactive requests and alert on runs that hit it.
  5. Add a test that streams a canned model response and asserts the frame sequence, and that the session contains the aggregate but no partials.
  6. If you need voice or interruptible sessions, prototype runLive with a LiveRequestQueue after confirming live-capable models in the current documentation.
Key takeaway: In ADK Java only StreamingMode.SSE turns on token streaming for runAsync. It produces partial events for each chunk and then one aggregated non-partial event. Tools run only from the aggregate, and the runner persists only non-partial events, so the session holds the settled text and never the chunks. Put a typed gateway between the Flowable and the browser, have clients append deltas and replace them with the aggregate, dispose subscriptions when clients leave, and use runLive with a LiveRequestQueue only when input and output must overlap.