Streaming changes what a user feels more than any other single setting on a Gemini-backed agent: the same answer that takes eight seconds to arrive feels slow as one block and fast when the first words appear after a few hundred milliseconds. ADK Java makes turning it on a one-line change, but the stream that reaches your code is shaped by what Gemini actually sends: text chunks, thought parts, a finish reason that can say SAFETY after half an answer has been shown, and token usage that only arrives at the end.
This article follows one streamed call from the Gemini wire to a browser, with the emphasis on the Gemini side: chunk anatomy, thoughts, finish reasons, usage and latency. How ADK's runtime turns chunks into events is covered in depth in streaming events in ADK Java and the Gemini class deep dive; here we use those mechanics and focus on what to do with them. API names below were checked against the adk-java and java-genai sources and the Gemini API reference in October 2026; both libraries move quickly, so re-check them against your versions.
What Gemini sends on the wire
Gemini exposes streaming as a separate method, models/{model}:streamGenerateContent. With alt=sse the response is a Server-Sent Events stream in which every data frame is a complete GenerateContentResponse JSON object, not a raw token. A chunk carries a candidate whose content holds one or more parts. Most chunks carry a slice of text; with thought summaries enabled, some parts are flagged thought: true; a function call appears as a part too. An abridged, illustrative final chunk looks like this:
data: {"candidates": [{"content": {"role": "model",
"parts": [{"text": " before the cut-off date."}]},
"finishReason": "STOP"}],
"usageMetadata": {"promptTokenCount": 812, "candidatesTokenCount": 164,
"thoughtsTokenCount": 233, "totalTokenCount": 1209},
"modelVersion": "..."}Three properties of this wire format drive everything else. Chunk boundaries are arbitrary, so a chunk may end mid-word or mid-markdown. The finish reason is meaningful only at the end: earlier chunks normally have none. And complete token counts, including thinking tokens, arrive with the last chunk, so cost is not final until the stream completes; ADK builds its final response from that last chunk. Google's newer guides also document a different, newer API surface with other field names; ADK Java's Gemini class calls streamGenerateContent, which is what this page describes. The documented finish reasons include STOP, MAX_TOKENS, SAFETY, RECITATION, LANGUAGE and OTHER, among others; a prompt rejected before generation instead returns promptFeedback.blockReason with no candidate at all.
From chunks to ADK events
Inside ADK Java, the Gemini model class calls the java-genai client's async generateContentStream and feeds the chunks through an aggregator. Each chunk with content becomes an LlmResponse marked partial. When the stream ends, the aggregator emits one non-partial response holding the accumulated parts, with thought text kept separate from answer text, and built from the last chunk, so it carries the finish reason and the usage metadata. In current adk-java, when that finish reason is anything other than STOP, the final response also gets errorCode and errorMessage set, while still carrying the content that was generated. The flow runs tools only from non-partial responses, and the runner persists only non-partial events. The detailed rules, including function-call id handling, are in LLM response streaming.
Turning streaming on
Streaming is selected per run through RunConfig. StreamingMode has three values: NONE (default), SSE for token streaming through runAsync, and BIDI for live sessions through runLive. A minimal console client that streams answer text, shows thoughts separately and reports how the turn ended:
GenerateContentConfig gen = GenerateContentConfig.builder()
.thinkingConfig(ThinkingConfig.builder().includeThoughts(true).build())
.build();
LlmAgent agent = LlmAgent.builder()
.name("explainer")
.model("gemini-2.5-flash")
.instruction("Answer concisely and cite the source document by title.")
.generateContentConfig(gen)
.build();
InMemoryRunner runner = new InMemoryRunner(agent, "demo");
RunConfig cfg = RunConfig.builder()
.streamingMode(RunConfig.StreamingMode.SSE)
.autoCreateSession(true)
.build();
runner.runAsync("u1", "s1", Content.fromParts(Part.fromText("Why is the sky blue?")), cfg)
.blockingForEach(ev -> {
boolean partial = ev.partial().orElse(false);
List<Part> parts = ev.content().flatMap(Content::parts).orElse(List.of());
if (partial) {
for (Part part : parts) {
String t = part.text().orElse("");
if (part.thought().orElse(false)) System.err.print(t); // thought summary
else System.out.print(t); // answer delta
}
return;
}
// non-partial: the settled turn; replace what was rendered, then inspect the ending
ev.errorCode().ifPresent(code ->
System.out.printf("%n[stopped: %s %s]%n", code, ev.errorMessage().orElse("")));
ev.usageMetadata().ifPresent(u -> System.out.printf("%n[tokens out=%s thoughts=%s]%n",
u.candidatesTokenCount().orElse(0), u.thoughtsTokenCount().orElse(0)));
});blockingForEach is fine for a demo; a server subscribes asynchronously and disposes the subscription when the client disconnects. The model id is an example; which models support which thinking options is set by Google and changes over time, so check model configuration and the current model list.
Thoughts in the stream
On thinking models the model reasons before it answers, and that reasoning costs time before the first answer token. Without includeThoughts, the stream is silent during that phase and the first visible chunk can arrive seconds later than on a non-thinking call. With it, Gemini streams thought summaries as parts flagged thought, so a UI can show progress in a collapsible reasoning panel. They are summaries, not the raw reasoning, and they are not answer text: never concatenate them into the reply, and never persist them as if they were.
Thinking is controlled through thinkingBudget or thinkingLevel on ThinkingConfig, depending on the model generation. A smaller budget or lower level cuts time to first answer token and output cost; a larger one helps multi-step problems. Thinking tokens are billed and reported separately as thoughtsTokenCount, which is only known at the end of the stream, so cost tracking must read the final event, never sum partials.
Endings that arrive after the text
The awkward property of streaming is that the verdict arrives last. A response can stream a paragraph and then end with SAFETY or RECITATION, or stop mid-sentence with MAX_TOKENS. The user has already seen the text. Decide per reason what the UI does, and make the final event authoritative:
| Ending | What the user saw | What to do |
|---|---|---|
STOP | Complete answer | Replace streamed text with the final event's text |
MAX_TOKENS | Truncated answer | Mark as cut off; offer continue; review maxOutputTokens |
SAFETY | Partial text, then a block | Retract or collapse streamed text; show a neutral notice |
RECITATION | Partial text | Retract; log for review; do not cache |
| Prompt blocked | Nothing | Error event with the block reason; no partials |
| Stream error | Partial text, then an error | Keep text marked incomplete; retry only with care |
Two implementation notes. First, read the outcome from the final non-partial event: in current adk-java its errorCode is set for any non-STOP ending, and finishReason() carries the raw value; older releases handled this differently, so test it on yours. Second, retrying a stream that failed after text was shown produces a different answer; either restart the bubble visibly or do not retry automatically.
Measuring a stream
Measure what users feel, not only total latency. Three numbers matter: time to first event of any kind, time to first answer token (after thoughts), and the largest gap between chunks, which shows up as stutter.
final class StreamClock {
final long start = System.nanoTime();
long firstAny = -1, firstAnswer = -1, last = -1, maxGap = 0; int chunks = 0;
void onEvent(Event ev) {
long now = System.nanoTime();
if (firstAny < 0) firstAny = now;
if (last >= 0) maxGap = Math.max(maxGap, now - last);
last = now;
if (ev.partial().orElse(false)) {
chunks++;
boolean answer = ev.content().flatMap(Content::parts).orElse(List.of()).stream()
.anyMatch(p -> !p.thought().orElse(false) && !p.text().orElse("").isEmpty());
if (answer && firstAnswer < 0) firstAnswer = now;
}
}
String report() {
return String.format("ttfe=%dms ttfa=%dms maxGap=%dms chunks=%d total=%dms",
ms(firstAny), ms(firstAnswer), maxGap / 1_000_000, chunks, ms(last));
}
private long ms(long t) { return t < 0 ? -1 : (t - start) / 1_000_000; }
}
// runner.runAsync(...).doOnNext(clock::onEvent).doOnComplete(() -> log.info(clock.report()))Export these as histograms per model and agent, alongside the final event's token counts. Time to first answer token is the number to watch when you change thinking settings; maximum gap is the one to watch when you change proxies or networks.
Getting the stream to the user, and when to use Live
Most broken streams are broken after ADK, between your server and the browser. A reverse proxy that buffers responses turns a stream back into one block; disable buffering for the streaming route. Response compression can also hold bytes until a buffer fills, so exclude text/event-stream from it or flush explicitly. Idle timeouts on load balancers cut long thinking phases, when no bytes flow; send a comment-line heartbeat every few seconds. Always dispose the ADK subscription when the client disconnects, or the model call keeps running and billing. A full gateway built on Spring's SseEmitter is in the streaming events article linked above.
Streaming text through runAsync is not the same as the Gemini Live API. Use runLive with a LiveRequestQueue when input and output must overlap, as in voice or interruptible sessions, and only with models that support live connections; its event flags and tool handling differ, as streaming tools in live mode shows. For a chat UI, SSE mode is simpler and sufficient.
Worked example: a slow-feeling assistant
A documentation assistant runs on a thinking model with includeThoughts off. Users complain that it feels slow, though median total latency is unchanged since a model upgrade. The stream clock shows the story: time to first event and time to first answer token are equal and long, because nothing is emitted while the model thinks, and the answer then arrives quickly. These figures are illustrative; your measurements will differ.
The team tries two fixes behind a flag. Enabling thought summaries makes the first event arrive early, and the UI shows a reasoning panel, but answer text is no sooner. Lowering the thinking budget for the documentation agent, whose questions are mostly lookups, cuts time to first answer token, and an offline evaluation on 200 saved questions shows no loss in answer correctness. They ship the lower budget, keep thought summaries for a separate troubleshooting agent where reasoning helps, and add an alert on the time-to-first-answer histogram. During rollout, the max-gap metric spikes in one region; the cause is a proxy with buffering re-enabled by a config change, caught because the metric existed.
Failure modes
| Failure | Cause | Fix |
|---|---|---|
| Whole answer appears at once | Proxy buffering or compression | Disable for the route; flush; check headers |
| Text duplicated in UI | Appending the final event to the deltas | Replace on the non-partial event |
| Reasoning shown as the answer | Thought parts not filtered | Branch on part.thought() |
| Cost underreported | Usage read from partials | Read usage from the final event only |
| Blocked text stays visible | Ignoring a non-STOP ending | Handle errorCode and finishReason on final |
| Stream cut during thinking | Idle timeout on a load balancer | Heartbeat comments; raise timeouts |
| Model keeps running after tab closes | Subscription never disposed | Dispose on client disconnect |
| Tools run on half a call | Acting on partial function calls | Let the flow run tools from the final |
Trade-offs
Streaming improves perceived latency but complicates everything downstream: clients must handle deltas and replacement, moderation must run on output that has already been shown, and retries become visible. For back-end callers, batch jobs and tools, leave StreamingMode.NONE and take one complete response. Thought summaries improve perceived progress and debuggability at the cost of extra output and UI complexity. Lower thinking settings buy speed at some risk to hard questions, so set them per agent from evaluation data, not globally.
What to do next
- Turn on
StreamingMode.SSEonly for user-facing agents; keep back-end calls non-streaming. - Render partial text as deltas and replace it with the final event's content.
- Branch on
part.thought()and decide whether thought summaries belong in your UI. - Handle every non-STOP ending explicitly, including retraction for SAFETY and RECITATION.
- Read token usage and thinking tokens from the final event only.
- Instrument time to first event, time to first answer token and maximum gap.
- Test the route through your real proxy and load balancer, and add a heartbeat.
- Dispose subscriptions on disconnect and verify billing stops.