ReAct, from the 2022 paper by Yao and colleagues, is the pattern behind most tool-using agents: the model reasons about what it knows, takes an action such as calling a tool, reads the result, and repeats until it can answer. The original paper did this in plain text, with the model writing lines like Thought, Action and Observation and a harness parsing them. Modern models do the same thing with structured function calling, and in the Agent Development Kit for Java that loop is built into LlmAgent.

That is both the good news and the trap. You get ReAct without writing a parser, but the loop's quality depends on things the framework cannot choose for you: what the instruction asks the model to reason about, how much reasoning it may spend, what observations look like, and when the loop must stop. This article maps each ReAct element onto ADK Java, builds an incident-triage agent, and adds the guards that keep it bounded. For the runtime stages under the loop, see the anatomy of the execution loop.

Advertisement

ReAct from first principles

A language model alone can only reason over what is in its context. A tool alone can only do what it is told. ReAct interleaves them: each reasoning step decides the next action, and each action's result becomes evidence for the next reasoning step. Compared with reasoning first and acting later, interleaving lets the agent correct course when a result surprises it, and grounding each step in a real observation reduces invented facts.

Three properties make it work. The model must be able to say which action it wants in a form the harness can execute unambiguously. Observations must come back into the context in a form the model can use. And something outside the model must end the loop, because a model that is unsure will keep acting.

How ADK Java implements each element

ReAct elementADK Java mechanismWhere you control it
ThoughtModel text before a call, and thought parts when thinking is enabledInstruction; ThinkingConfig in generateContentConfig
ActionA FunctionCall part in a model eventTool declarations: names, descriptions, parameter schemas
ObservationA FunctionResponse event appended to the sessionWhat the tool method returns
LoopThe LLM flow calls the model again while the last response contained callsRunConfig, callbacks
StopA response without function calls is finalInstruction, tool budget, maxLlmCalls
ReAct mapped onto one ADK Java invocationUser messagerunner.runAsyncLLM call (Thought)thought parts + textFunctionCall (Action)name + args + idbeforeToolCallbackbudget, dedupe, policyTool runsFunctionTool methodFunctionResponseObservation eventSession historyall prior eventsFinal answerfinalResponse() truetool callallowedappendnext callno callEvery trip round the loop costs one model call and counts against RunConfig.maxLlmCalls (default 500).A callback that returns a map skips the tool and hands that map to the model as the observation.
One invocation: each model call may emit function calls; allowed calls run and their responses are appended to the session; the next call sees them. A response without calls ends the loop.

Because actions are structured function calls, there is nothing to parse and no risk of the harness misreading free text. The observation for each call is matched back by call id, which is how parallel calls in one response stay paired with their results; the declaration side is covered in Gemini function calling in ADK Java.

One note on planners. The Python ADK ships a PlanReActPlanner that makes a model without native thinking write explicit plan and reasoning sections. When this was written, the Java API reference did not list it: the Java com.google.adk.planner package contains multi-agent orchestrators (sequential, parallel, loop and supervisor planners), not a ReAct prompt planner. In Java you get the same effect with the instruction and with model thinking, as below.

Advertisement

Worked example: an incident-triage agent

The agent answers questions like "checkout errors spiked at 14:05, why?". It has three read-only tools. Each returns a small map with an explicit status, so failures come back as observations the model can reason about instead of exceptions that end the invocation.

public class TriageTools {

  @Schema(description = "Recent deploys of a service, newest first. Returns at most 5.")
  public static Map<String, Object> recentDeploys(
      @Schema(name = "service", description = "Service name, e.g. checkout") String service) {
    List<Deploy> ds = DeployStore.latest(service, 5);
    return Map.of("status", "ok", "deploys", ds.stream().map(Deploy::summary).toList());
  }

  @Schema(description = "Error-rate percentage for a service over the last N minutes (5-120).")
  public static Map<String, Object> errorRate(
      @Schema(name = "service", description = "Service name") String service,
      @Schema(name = "minutes", description = "Window in minutes, 5 to 120") int minutes) {
    if (minutes < 5 || minutes > 120) {
      return Map.of("status", "error", "message", "minutes must be 5-120");   // observation, not exception
    }
    return Map.of("status", "ok", "errorPct", Metrics.errorPct(service, minutes));
  }

  @Schema(description = "Up to 20 matching log lines, truncated to 200 characters each.")
  public static Map<String, Object> searchLogs(
      @Schema(name = "service", description = "Service name") String service,
      @Schema(name = "query", description = "Plain-text search") String query) {
    return Map.of("status", "ok", "lines", LogIndex.search(service, query, 20, 200));
  }
}

The instruction does the ReAct work the paper did with prompt templates: it asks for a stated expectation before each call and an update after each result, forbids repeated calls, and defines when to stop and what the answer must contain. Temperature 0 makes the trajectory repeatable enough to test.

String instruction = """
    You triage production incidents. Work in steps.
    Before each tool call, state in one sentence what you expect to learn.
    After each result, state what it changed about your hypothesis.
    Never call the same tool with the same arguments twice.
    Stop and answer when one cause explains the evidence, or after 6 tool calls
    say what you checked and what remains unknown.
    Final answer: cause, evidence (tool results you relied on), suggested next action.
    """;

LlmAgent triage = LlmAgent.builder()
    .name("incident_triage")
    .model(System.getenv("ADK_MODEL"))
    .instruction(instruction)
    .tools(FunctionTool.create(TriageTools.class, "recentDeploys"),
           FunctionTool.create(TriageTools.class, "errorRate"),
           FunctionTool.create(TriageTools.class, "searchLogs"))
    .generateContentConfig(GenerateContentConfig.builder()
        .temperature(0.0f)
        .thinkingConfig(ThinkingConfig.builder()
            .includeThoughts(true)          // return thought summaries, if the model supports it
            .thinkingBudget(1024)           // tokens; 0 disables, -1 automatic; ranges vary by model
            .build())
        .build())
    .beforeToolCallback(ReactGuards::check)
    .build();

RunConfig cfg = RunConfig.builder().maxLlmCalls(12).build();   // hard ceiling per invocation

A good trajectory looks like this. The model states that a recent deploy is the likeliest cause and calls recentDeploys("checkout"). The observation shows a deploy at 14:02. It calls errorRate("checkout", 30) to confirm the spike, then searchLogs("checkout", "payment timeout") and sees timeouts to a payment client changed in that deploy. One cause explains the evidence, so it answers with the deploy, the three observations and a rollback suggestion. That is four model calls: three that emit actions and one final answer.

Thoughts: enable, capture, keep internal

With includeThoughts(true), models that support thinking return thought summaries as parts marked thought. They are the most useful debugging signal you have, because they show why the agent chose an action. They are not an answer: they can contain half-formed hypotheses, internal tool names, or data from earlier observations. Log them to a restricted trace, not to the user.

runner.runAsync(userId, session.id(), msg, cfg).blockingForEach(ev -> {
  ev.content().flatMap(Content::parts).orElse(List.of()).forEach(part -> {
    if (part.thought().orElse(false)) {
      traceLog.debug("THOUGHT {}", part.text().orElse(""));     // internal trace only
    }
  });
  ev.functionCalls().forEach(fc ->
      traceLog.info("ACTION {} {}", fc.name().orElse("?"), fc.args().orElse(Map.of())));
  ev.functionResponses().forEach(fr ->
      traceLog.info("OBSERVATION {}", fr.name().orElse("?")));
  if (ev.finalResponse()) {
    reply(ev.stringifyContent());
  }
});

The thinking budget is a cost and latency lever. A larger budget helps on multi-hop diagnosis; for a simple lookup agent it only adds tokens. Measure trajectory quality at two or three budgets on your own evaluation set rather than guessing, since allowed ranges and defaults differ by model.

Bounding the loop

ReAct agents fail open: an unsure model keeps acting. ADK Java gives you two layers of limits. RunConfig.maxLlmCalls is the hard ceiling on model calls per invocation; the default is 500, far too high for an interactive agent, and exceeding it raises LlmCallsLimitExceededException down the stream, which your caller must turn into a useful reply. The runtime is described in runtime configuration in depth.

The softer layer is a beforeToolCallback. Returning a map skips the tool and gives that map to the model as the observation; returning empty runs the tool. That lets you enforce a tool budget and reject duplicate calls while still giving the model a chance to answer gracefully, instead of killing the invocation.

final class ReactGuards {
  static final int MAX_TOOL_CALLS = 6;

  record Budget(AtomicInteger calls, Set<String> seen, Instant started) {}
  private static final ConcurrentHashMap<String, Budget> BY_INVOCATION = new ConcurrentHashMap<>();

  static Maybe<Map<String, Object>> check(InvocationContext ctx, BaseTool tool,
                                          Map<String, Object> args, ToolContext tc) {
    Budget b = BY_INVOCATION.computeIfAbsent(ctx.invocationId(),
        id -> new Budget(new AtomicInteger(), ConcurrentHashMap.newKeySet(), Instant.now()));
    if (!b.seen().add(tool.name() + ":" + new TreeMap<>(args))) {     // atomic check-and-add
      return Maybe.just(Map.of("status", "error",
          "message", "Duplicate call; you already have this result. Use it or try something else."));
    }
    if (b.calls().incrementAndGet() > MAX_TOOL_CALLS) {              // atomic, safe for parallel calls
      return Maybe.just(Map.of("status", "error",
          "message", "Tool budget exhausted. Answer with what you have."));
    }
    return Maybe.empty();                       // empty = run the real tool
  }

  /** Call from a scheduler: drop budgets of invocations that must have finished. */
  static void evictOlderThan(Duration age) {
    Instant cutoff = Instant.now().minus(age);
    BY_INVOCATION.values().removeIf(b -> b.started().isBefore(cutoff));
  }
}

The budget is keyed by invocationId(), so it resets on every user turn, and both checks are atomic because calls from one model response can run in parallel. Session state would be the wrong home: State.put writes through to the live session map, so a counter kept there can outlive the turn. The key sorts the arguments so reordered arguments still match. Set maxLlmCalls a little above the tool budget plus one final answer, so the soft limit fires first.

Designing observations

The observation is the only way the world reaches the model's reasoning, so its shape matters as much as the prompt. Keep observations small: cap list lengths and truncate text in the tool, since every observation stays in the context for all later calls. Make them explicit: a status field, units in field names or descriptions, and absolute timestamps. Return errors as data, with a message that says what to change, because a model can correct a bad argument but cannot recover from an exception that aborts the run.

Treat observations as untrusted input. Log lines, tickets and web pages can contain text that reads like instructions, and a ReAct agent will reason over it. Keep write-capable tools out of agents that read untrusted content, or put a confirmation step in front of them.

Failure modes

SymptomCauseFix
Same tool called again and againObservation did not answer the question, or was too vagueDuplicate guard; richer, explicit observation
Invocation dies with a limit exceptionmaxLlmCalls hit before the model answeredSoft tool budget below the hard limit; catch and reply
Answer ignores a tool resultResult buried in a large observationTruncate and summarise inside the tool
Agent answers without actingInstruction does not require evidenceRequire cited observations in the final answer
Invented causeModel filled a gap instead of stoppingAllow and ask for an explicit unknown
Costs creep upThinking budget and context both grow per stepMeasure per-step tokens; lower budget; cap observations

Operating and evaluating a ReAct agent

Evaluate trajectories, not just answers. For each test case record the sequence of actions and check it against expectations: the required tools were called, forbidden ones were not, and the step count is within budget. A correct answer reached in eleven steps is a cost bug waiting to happen. Replay a fixed set of incidents on every prompt or model change and diff the trajectories.

In production, emit one span per model call and per tool call with the step number, and alert on the distribution of steps per invocation and on guard rejections, which are early signs of a confused agent. Tool observability and metrics covers the built-in spans. The main trade-off is depth against cost: every extra step adds a model call carrying the whole growing context, so the second half of a long trajectory costs more than the first.

What to do next

  1. Write the instruction as a ReAct contract: expectation before each call, update after each result, a stop rule and an answer format that cites evidence.
  2. Return status-bearing maps from every tool; make errors observations, not exceptions.
  3. Cap list sizes and text length inside each tool.
  4. Add a beforeToolCallback with a per-invocation tool budget and duplicate-call guard, keyed by invocationId.
  5. Set maxLlmCalls just above the tool budget and handle LlmCallsLimitExceededException in the caller.
  6. Enable thought summaries on a supported model and log them to a restricted trace.
  7. Build a trajectory test set and compare step counts and tool sequences on every change.
Key takeaway: ADK Java runs ReAct natively: model calls are the thoughts, FunctionCall parts are the actions, FunctionResponse events are the observations, and the loop ends at a response without calls. What the framework cannot decide is quality and bounds. Write the instruction as a reasoning contract with a stop rule, return small explicit observations with errors as data, enable and log thoughts internally, enforce a soft tool budget and duplicate guard in beforeToolCallback, keep maxLlmCalls as the hard ceiling, and evaluate whole trajectories.