Streaming changes what a user feels more than any other single setting on a Gemini-backed agent: the same answer that takes eight seconds to arrive feels slow as one block and fast when the first words appear after a few hundred milliseconds. ADK Java makes turning it on a one-line change, but the stream that reaches your code is shaped by what Gemini actually sends: text chunks, thought parts, a finish reason that can say SAFETY after half an answer has been shown, and token usage that only arrives at the end.

This article follows one streamed call from the Gemini wire to a browser, with the emphasis on the Gemini side: chunk anatomy, thoughts, finish reasons, usage and latency. How ADK's runtime turns chunks into events is covered in depth in streaming events in ADK Java and the Gemini class deep dive; here we use those mechanics and focus on what to do with them. API names below were checked against the adk-java and java-genai sources and the Gemini API reference in October 2026; both libraries move quickly, so re-check them against your versions.

What Gemini sends on the wire

Gemini exposes streaming as a separate method, models/{model}:streamGenerateContent. With alt=sse the response is a Server-Sent Events stream in which every data frame is a complete GenerateContentResponse JSON object, not a raw token. A chunk carries a candidate whose content holds one or more parts. Most chunks carry a slice of text; with thought summaries enabled, some parts are flagged thought: true; a function call appears as a part too. An abridged, illustrative final chunk looks like this:

data: {"candidates": [{"content": {"role": "model",
        "parts": [{"text": " before the cut-off date."}]},
        "finishReason": "STOP"}],
       "usageMetadata": {"promptTokenCount": 812, "candidatesTokenCount": 164,
        "thoughtsTokenCount": 233, "totalTokenCount": 1209},
       "modelVersion": "..."}

Three properties of this wire format drive everything else. Chunk boundaries are arbitrary, so a chunk may end mid-word or mid-markdown. The finish reason is meaningful only at the end: earlier chunks normally have none. And complete token counts, including thinking tokens, arrive with the last chunk, so cost is not final until the stream completes; ADK builds its final response from that last chunk. Google's newer guides also document a different, newer API surface with other field names; ADK Java's Gemini class calls streamGenerateContent, which is what this page describes. The documented finish reasons include STOP, MAX_TOKENS, SAFETY, RECITATION, LANGUAGE and OTHER, among others; a prompt rejected before generation instead returns promptFeedback.blockReason with no candidate at all.

From chunks to ADK events

One streamed Gemini call through ADK Java (SSE mode)Gemini API:streamGenerateContentjava-genai clientasync generateContentStreamAggregatorin Gemini.javaBaseLlmFlowtools on final onlySSE chunksPartial LlmResponseper content chunkFinal LlmResponseparts + reason + usageRunner: Flowable of Event to your subscriberpartials emitted, only non-partial events persistedYour gateway and browserrender deltas, replace on final, show reasonChunk n: text or thought partsLast chunk: finishReason,usageMetadata
Gemini sends complete response objects as SSE frames. ADK's aggregator re-emits each content chunk as a partial response, then one final response with the accumulated parts, the last chunk's finish reason and its usage.

Inside ADK Java, the Gemini model class calls the java-genai client's async generateContentStream and feeds the chunks through an aggregator. Each chunk with content becomes an LlmResponse marked partial. When the stream ends, the aggregator emits one non-partial response holding the accumulated parts, with thought text kept separate from answer text, and built from the last chunk, so it carries the finish reason and the usage metadata. In current adk-java, when that finish reason is anything other than STOP, the final response also gets errorCode and errorMessage set, while still carrying the content that was generated. The flow runs tools only from non-partial responses, and the runner persists only non-partial events. The detailed rules, including function-call id handling, are in LLM response streaming.

Turning streaming on

Streaming is selected per run through RunConfig. StreamingMode has three values: NONE (default), SSE for token streaming through runAsync, and BIDI for live sessions through runLive. A minimal console client that streams answer text, shows thoughts separately and reports how the turn ended:

GenerateContentConfig gen = GenerateContentConfig.builder()
    .thinkingConfig(ThinkingConfig.builder().includeThoughts(true).build())
    .build();
LlmAgent agent = LlmAgent.builder()
    .name("explainer")
    .model("gemini-2.5-flash")
    .instruction("Answer concisely and cite the source document by title.")
    .generateContentConfig(gen)
    .build();

InMemoryRunner runner = new InMemoryRunner(agent, "demo");
RunConfig cfg = RunConfig.builder()
    .streamingMode(RunConfig.StreamingMode.SSE)
    .autoCreateSession(true)
    .build();

runner.runAsync("u1", "s1", Content.fromParts(Part.fromText("Why is the sky blue?")), cfg)
    .blockingForEach(ev -> {
      boolean partial = ev.partial().orElse(false);
      List<Part> parts = ev.content().flatMap(Content::parts).orElse(List.of());
      if (partial) {
        for (Part part : parts) {
          String t = part.text().orElse("");
          if (part.thought().orElse(false)) System.err.print(t);   // thought summary
          else System.out.print(t);                                 // answer delta
        }
        return;
      }
      // non-partial: the settled turn; replace what was rendered, then inspect the ending
      ev.errorCode().ifPresent(code ->
          System.out.printf("%n[stopped: %s %s]%n", code, ev.errorMessage().orElse("")));
      ev.usageMetadata().ifPresent(u -> System.out.printf("%n[tokens out=%s thoughts=%s]%n",
          u.candidatesTokenCount().orElse(0), u.thoughtsTokenCount().orElse(0)));
    });

blockingForEach is fine for a demo; a server subscribes asynchronously and disposes the subscription when the client disconnects. The model id is an example; which models support which thinking options is set by Google and changes over time, so check model configuration and the current model list.

Thoughts in the stream

On thinking models the model reasons before it answers, and that reasoning costs time before the first answer token. Without includeThoughts, the stream is silent during that phase and the first visible chunk can arrive seconds later than on a non-thinking call. With it, Gemini streams thought summaries as parts flagged thought, so a UI can show progress in a collapsible reasoning panel. They are summaries, not the raw reasoning, and they are not answer text: never concatenate them into the reply, and never persist them as if they were.

Thinking is controlled through thinkingBudget or thinkingLevel on ThinkingConfig, depending on the model generation. A smaller budget or lower level cuts time to first answer token and output cost; a larger one helps multi-step problems. Thinking tokens are billed and reported separately as thoughtsTokenCount, which is only known at the end of the stream, so cost tracking must read the final event, never sum partials.

Endings that arrive after the text

The awkward property of streaming is that the verdict arrives last. A response can stream a paragraph and then end with SAFETY or RECITATION, or stop mid-sentence with MAX_TOKENS. The user has already seen the text. Decide per reason what the UI does, and make the final event authoritative:

EndingWhat the user sawWhat to do
STOPComplete answerReplace streamed text with the final event's text
MAX_TOKENSTruncated answerMark as cut off; offer continue; review maxOutputTokens
SAFETYPartial text, then a blockRetract or collapse streamed text; show a neutral notice
RECITATIONPartial textRetract; log for review; do not cache
Prompt blockedNothingError event with the block reason; no partials
Stream errorPartial text, then an errorKeep text marked incomplete; retry only with care

Two implementation notes. First, read the outcome from the final non-partial event: in current adk-java its errorCode is set for any non-STOP ending, and finishReason() carries the raw value; older releases handled this differently, so test it on yours. Second, retrying a stream that failed after text was shown produces a different answer; either restart the bubble visibly or do not retry automatically.

Measuring a stream

Measure what users feel, not only total latency. Three numbers matter: time to first event of any kind, time to first answer token (after thoughts), and the largest gap between chunks, which shows up as stutter.

final class StreamClock {
  final long start = System.nanoTime();
  long firstAny = -1, firstAnswer = -1, last = -1, maxGap = 0; int chunks = 0;

  void onEvent(Event ev) {
    long now = System.nanoTime();
    if (firstAny < 0) firstAny = now;
    if (last >= 0) maxGap = Math.max(maxGap, now - last);
    last = now;
    if (ev.partial().orElse(false)) {
      chunks++;
      boolean answer = ev.content().flatMap(Content::parts).orElse(List.of()).stream()
          .anyMatch(p -> !p.thought().orElse(false) && !p.text().orElse("").isEmpty());
      if (answer && firstAnswer < 0) firstAnswer = now;
    }
  }
  String report() {
    return String.format("ttfe=%dms ttfa=%dms maxGap=%dms chunks=%d total=%dms",
        ms(firstAny), ms(firstAnswer), maxGap / 1_000_000, chunks, ms(last));
  }
  private long ms(long t) { return t < 0 ? -1 : (t - start) / 1_000_000; }
}
// runner.runAsync(...).doOnNext(clock::onEvent).doOnComplete(() -> log.info(clock.report()))

Export these as histograms per model and agent, alongside the final event's token counts. Time to first answer token is the number to watch when you change thinking settings; maximum gap is the one to watch when you change proxies or networks.

Getting the stream to the user, and when to use Live

Most broken streams are broken after ADK, between your server and the browser. A reverse proxy that buffers responses turns a stream back into one block; disable buffering for the streaming route. Response compression can also hold bytes until a buffer fills, so exclude text/event-stream from it or flush explicitly. Idle timeouts on load balancers cut long thinking phases, when no bytes flow; send a comment-line heartbeat every few seconds. Always dispose the ADK subscription when the client disconnects, or the model call keeps running and billing. A full gateway built on Spring's SseEmitter is in the streaming events article linked above.

Streaming text through runAsync is not the same as the Gemini Live API. Use runLive with a LiveRequestQueue when input and output must overlap, as in voice or interruptible sessions, and only with models that support live connections; its event flags and tool handling differ, as streaming tools in live mode shows. For a chat UI, SSE mode is simpler and sufficient.

Worked example: a slow-feeling assistant

A documentation assistant runs on a thinking model with includeThoughts off. Users complain that it feels slow, though median total latency is unchanged since a model upgrade. The stream clock shows the story: time to first event and time to first answer token are equal and long, because nothing is emitted while the model thinks, and the answer then arrives quickly. These figures are illustrative; your measurements will differ.

The team tries two fixes behind a flag. Enabling thought summaries makes the first event arrive early, and the UI shows a reasoning panel, but answer text is no sooner. Lowering the thinking budget for the documentation agent, whose questions are mostly lookups, cuts time to first answer token, and an offline evaluation on 200 saved questions shows no loss in answer correctness. They ship the lower budget, keep thought summaries for a separate troubleshooting agent where reasoning helps, and add an alert on the time-to-first-answer histogram. During rollout, the max-gap metric spikes in one region; the cause is a proxy with buffering re-enabled by a config change, caught because the metric existed.

Failure modes

FailureCauseFix
Whole answer appears at onceProxy buffering or compressionDisable for the route; flush; check headers
Text duplicated in UIAppending the final event to the deltasReplace on the non-partial event
Reasoning shown as the answerThought parts not filteredBranch on part.thought()
Cost underreportedUsage read from partialsRead usage from the final event only
Blocked text stays visibleIgnoring a non-STOP endingHandle errorCode and finishReason on final
Stream cut during thinkingIdle timeout on a load balancerHeartbeat comments; raise timeouts
Model keeps running after tab closesSubscription never disposedDispose on client disconnect
Tools run on half a callActing on partial function callsLet the flow run tools from the final

Trade-offs

Streaming improves perceived latency but complicates everything downstream: clients must handle deltas and replacement, moderation must run on output that has already been shown, and retries become visible. For back-end callers, batch jobs and tools, leave StreamingMode.NONE and take one complete response. Thought summaries improve perceived progress and debuggability at the cost of extra output and UI complexity. Lower thinking settings buy speed at some risk to hard questions, so set them per agent from evaluation data, not globally.

What to do next

  1. Turn on StreamingMode.SSE only for user-facing agents; keep back-end calls non-streaming.
  2. Render partial text as deltas and replace it with the final event's content.
  3. Branch on part.thought() and decide whether thought summaries belong in your UI.
  4. Handle every non-STOP ending explicitly, including retraction for SAFETY and RECITATION.
  5. Read token usage and thinking tokens from the final event only.
  6. Instrument time to first event, time to first answer token and maximum gap.
  7. Test the route through your real proxy and load balancer, and add a heartbeat.
  8. Dispose subscriptions on disconnect and verify billing stops.
Key takeaway: Gemini streams complete response objects whose text and thought parts arrive in arbitrary slices, with the finish reason and token usage only at the end. In ADK Java, SSE mode turns those chunks into partial events plus one final event that carries the settled content, the ending and the usage. Render partials, trust only the final event, handle non-STOP endings, and measure time to first answer token.