Temperature is the one sampling knob almost every agent developer touches, and the one most often set by folklore: zero for anything serious, high for anything creative. In ADK Java it is a field on the GenerateContentConfig you attach to an LlmAgent, so setting it takes one line. Choosing it well, applying it per call when one agent does several kinds of work, and knowing what it cannot guarantee takes more than that.
This article starts from what temperature does to the model's next-token distribution, with the numbers worked out, then shows exactly where it lives in ADK Java: the agent builder, a beforeModelCallback that overrides it per request, and logging in an afterModelCallback. It covers why temperature zero is not determinism, why newer Gemini models want you to leave it alone, how to measure the effect instead of guessing, and the failure modes that come from getting it wrong. API names are from the adk-java source and the Google Gen AI Java SDK; check them against the version you build with.
What temperature does to the next token
At each step the model produces a logit z_i for every token in its vocabulary. Sampling turns logits into probabilities with a softmax, and temperature T divides the logits first: p_i = exp(z_i / T) divided by the sum over j of exp(z_j / T). T below 1 sharpens the distribution toward the top token, T above 1 flattens it, and as T approaches 0 the sampler always takes the most likely token, which is greedy decoding.
Work it for three candidate tokens with logits 2.0, 1.0 and 0.1:
| Temperature | p(token A) | p(token B) | p(token C) |
|---|---|---|---|
| 0.5 | 0.864 | 0.117 | 0.019 |
| 1.0 | 0.659 | 0.242 | 0.099 |
| 2.0 | 0.502 | 0.304 | 0.194 |
At T = 1 the model picks the runner-up about a quarter of the time; at T = 0.5 about one time in nine; at T = 2 the third choice, which the model rated far lower, appears almost one time in five. Over a 500-token answer those per-token odds compound, which is why a high temperature changes not just word choice but the path the answer takes, and why a low temperature makes repeated runs look alike.
Two neighbours interact with temperature. Top-p (nucleus) sampling keeps the smallest set of tokens whose probabilities sum to p and samples among them; top-k keeps the k most likely. Both cut the tail that a high temperature fattens. Change one knob at a time: tuning temperature and top-p together makes it impossible to tell which one fixed or broke an output.
Setting temperature on an LlmAgent
In ADK Java the knob lives on com.google.genai.types.GenerateContentConfig from the Google Gen AI Java SDK, passed to the agent builder. The agent's config becomes the config of every LlmRequest it sends:
import com.google.adk.agents.LlmAgent;
import com.google.genai.types.GenerateContentConfig;
LlmAgent classifier = LlmAgent.builder()
.name("ticket_classifier")
.model("gemini-2.5-flash")
.instruction("Classify the support ticket into exactly one of: billing, outage, account, other.")
.generateContentConfig(GenerateContentConfig.builder()
.temperature(0.0f)
.maxOutputTokens(256)
.build())
.build();
LlmAgent copywriter = LlmAgent.builder()
.name("release_note_writer")
.model("gemini-2.5-flash")
.instruction("Write three alternative one-line release-note headlines for the change described.")
.generateContentConfig(GenerateContentConfig.builder()
.temperature(1.0f)
.topP(0.95f)
.build())
.build();The setting belongs to the agent, not the session or the runner. In a multi-agent tree each sub-agent has its own config, so a low-temperature router can delegate to a high-temperature writer without either affecting the other, and a sub-agent that sets nothing gets the model's default, not its parent's value. Keep tools, system instructions and output schemas on the agent builder rather than inside this config: the agent and its flow assemble those parts of the request themselves.
Note the types: the Gen AI SDK's builder takes a Float for temperature and top-p, so write 0.2f, not 0.2. A practical pattern is to build the config from a small record loaded from configuration, so that every agent's sampling settings live in one reviewed file instead of being scattered through builder chains:
record Sampling(float temperature, float topP, int maxOutputTokens) {
GenerateContentConfig toConfig() {
return GenerateContentConfig.builder()
.temperature(temperature)
.topP(topP)
.maxOutputTokens(maxOutputTokens)
.build();
}
}
// e.g. loaded from application config keyed by agent name
Sampling s = samplingFor("ticket_classifier");
LlmAgent agent = LlmAgent.builder()
.name("ticket_classifier")
.model("gemini-2.5-flash")
.instruction("...")
.generateContentConfig(s.toConfig())
.build();That also makes the experiment later in this article a configuration change rather than a code change.
One caution on the examples: gemini-2.5-flash is a thinking model, and its internal reasoning is sampled too. Temperature zero on such a model is a reasonable starting point for a narrow classifier, but treat it as a hypothesis for the experiment below rather than a rule, and do not set the output-token cap so tight that a thinking model runs out of room before it answers. If the response comes back empty or cut off, raise maxOutputTokens before touching temperature.
Top-p and top-k deserve the same discipline. On a classifier with a fixed answer set they rarely matter, because almost all probability sits on a handful of tokens. On a generator they decide how much of the flattened tail a high temperature can reach: temperature 1.0 with top-p 0.95 gives variety while still trimming the improbable tokens that produce non sequiturs. Leave top-k unset unless you have measured a reason to use it, since a fixed count of candidates behaves differently at every position in the text.
Per-call overrides with a before-model callback
One agent sometimes does two kinds of work: a planner that brainstorms options and then commits to one, or an agent whose behaviour should depend on a flag in session state. ADK Java's beforeModelCallback receives the LlmRequest.Builder before the model is called, so it can replace the config for that call only. The synchronous form, registered with beforeModelCallbackSync, returns Optional<LlmResponse>; returning empty lets the call proceed. The asynchronous form, registered with beforeModelCallback, returns an RxJava Maybe<LlmResponse> instead.
import com.google.adk.agents.Callbacks.BeforeModelCallbackSync;
import com.google.genai.types.GenerateContentConfig;
import java.util.Optional;
BeforeModelCallbackSync temperatureFromState = (ctx, request) -> {
Object mode = ctx.state().get("sampling_mode"); // set by an earlier step or by the caller
float t = "explore".equals(mode) ? 1.0f : 0.2f;
GenerateContentConfig base = request.config()
.orElseGet(() -> GenerateContentConfig.builder().build());
request.config(base.toBuilder().temperature(t).build()); // keep every other field
return Optional.empty(); // continue to the model
};
LlmAgent planner = LlmAgent.builder()
.name("planner")
.model("gemini-2.5-flash")
.instruction("...")
.beforeModelCallbackSync(temperatureFromState)
.build();Two details matter. Start from the existing config with toBuilder() instead of building a fresh one, or you silently drop maxOutputTokens, safety settings and anything the framework added. And keep the decision visible: write the chosen temperature into an afterModelCallback log line or a trace attribute along with the agent name, so that when an output looks wrong you can see which setting produced it. The callback ordering and request assembly are covered in model call orchestration in the ADK Java runtime.
Temperature zero is not determinism
Temperature zero is the most common setting in agent code and the most misunderstood. It makes sampling greedy, so each step takes the top token. It does not make the whole call reproducible. Serving systems batch requests together, and the floating-point sums inside a forward pass can come out in a different order depending on batch composition and hardware, which nudges logits by tiny amounts. When two tokens are nearly tied, a tiny nudge flips the choice, and every later token follows from that flip. Google's documentation describes low temperature as mostly, not fully, deterministic for this reason.
If you need repeatability for tests, the config also has a seed field, which asks the service to sample reproducibly when everything else is identical; treat it as best effort, not a contract. The reliable tools for repeatability are structural: constrain the output with an output schema, an enum or a short answer space; validate and retry; and in tests, compare outputs semantically rather than byte for byte. The LLM-as-judge scorer article uses exactly this combination for a judge agent: temperature zero plus an output schema.
There is a second reason not to reach for zero by reflex. Google's Gemini 3 developer guidance recommends keeping temperature at its default of 1.0 and warns that lowering it can cause looping or degraded performance, especially on reasoning-heavy tasks. Reasoning models are tuned to sample their thinking at the default temperature. So the right value depends on the model: earlier Gemini generations generally behave as the folklore says, newer reasoning models may not. Check the guidance for the exact model in your .model(...) call, and see the Gemini variant comparison for how variants differ in ADK Java.
Measuring instead of guessing
Choose a temperature with an experiment, not a guess. The harness is small: run the same input N times at each candidate temperature and score two things, correctness against expected answers and diversity across runs.
import com.google.adk.runner.InMemoryRunner;
import com.google.adk.sessions.Session;
import com.google.genai.types.Content;
import com.google.genai.types.Part;
import java.util.*;
static String runOnce(LlmAgent agent, String input) {
InMemoryRunner runner = new InMemoryRunner(agent);
Session session = runner.sessionService()
.createSession(runner.appName(), "eval-user").blockingGet();
Content msg = Content.fromParts(Part.fromText(input));
StringBuilder out = new StringBuilder();
runner.runAsync("eval-user", session.id(), msg).blockingForEach(ev -> {
if (ev.finalResponse()) out.append(ev.stringifyContent());
});
return out.toString().trim();
}
for (float t : new float[] {0.0f, 0.4f, 0.8f, 1.0f}) {
LlmAgent agent = classifierWithTemperature(t); // same instruction, only t differs
Map<String, Integer> distinct = new HashMap<>();
int correct = 0;
for (int i = 0; i < 20; i++) {
String answer = runOnce(agent, TICKET);
distinct.merge(answer, 1, Integer::sum);
if (answer.equalsIgnoreCase(EXPECTED)) correct++;
}
System.out.printf("T=%.1f correct=%d/20 distinct=%d%n", t, correct, distinct.size());
}Run it over a set of inputs, not one, and include hard cases. For a classifier you want the lowest temperature at which accuracy holds and distinct-answer counts stay at one per input. For a generator you want enough diversity that three candidates are actually different, and you want a quality score from a rubric or a judge so diversity is not bought with nonsense. Twenty samples per point will not detect small differences; use it to find gross effects, then confirm with a larger evaluation as described in the ADK Java evaluation overview.
| Task | Starting point | Why |
|---|---|---|
| Routing, classification, extraction to a schema | Low (0 to 0.3) on non-reasoning models | One right answer; variance is pure risk |
| Tool-calling agents | Low to moderate | Argument errors grow with randomness; keep diversity in the plan, not the arguments |
| Drafting, brainstorming, candidate generation | Default to 1.0 | You want different samples to pick from |
| Reasoning models such as Gemini 3 | Model default | Vendor guidance warns against lowering it |
Failure modes and trade-offs
What goes wrong in practice:
- Repetition loops at low temperature. Greedy decoding can fall into a cycle, repeating a phrase or re-calling the same tool. The fix is a higher temperature, a frequency or presence penalty where the model supports it, or a step limit in the agent loop; not a longer instruction.
- Hallucinated tool arguments at high temperature. A copywriter-level temperature on an agent that calls tools produces invented IDs and malformed JSON. Split the work: a creative agent proposes, a low-temperature agent executes.
- Flaky tests blamed on temperature. Setting zero and asserting exact strings still fails occasionally, for the batching reason above. Assert on structure and meaning.
- Silently lost settings. Building a new config in a callback instead of using
toBuilder()drops other fields; a sub-agent missing a config runs at the model default regardless of what its parent uses. - Configuration drift. Temperature hard-coded in twenty agents cannot be tuned. Read it from configuration per agent, as in runtime configuration for ADK Java, and record it in traces.
The trade-off underneath all of these is the same: lower temperature buys consistency and costs exploration, and the price of each depends on the model and the task. Structure (schemas, enums, validation, retries) buys consistency far more reliably than temperature does, so lean on structure first and use temperature to tune what is left.
What to do next
- List every LlmAgent in your application with its model and temperature; flag any that set none.
- Move temperature values into configuration keyed by agent name, and record the value used in an afterModelCallback log line or trace attribute.
- For agents that need two modes, add a beforeModelCallback that derives temperature from session state using
toBuilder(). - Check the vendor guidance for each model in use; leave reasoning models such as Gemini 3 at their default unless an experiment says otherwise.
- Run the 20-sample harness at three or four temperatures on a representative input set, scoring accuracy and distinct outputs.
- Replace exact-string assertions in tests with schema checks or semantic comparisons, and add output schemas to agents whose answers have a fixed shape.