An agent hallucinates when its answer states something that none of its inputs support: not the user's message, not the instruction, not any tool result in the turn. In a chat demo that is an embarrassment. In an agent that quotes order status, policy terms or account balances, it is an incident. Teams usually react by adding a sentence to the prompt and moving on, and the problem returns in a different shape a week later, because nobody found out which layer produced it.

This article treats a hallucination as a bug with a root cause and gives you a workflow to find it in an ADK Java agent: record exactly what the model saw and said, replay the turn with recorded tool results until the failure reproduces at a measurable rate, classify the cause with a fixed decision procedure, and fix it at the layer that produced it. The code targets google-adk 1.10.1; callback signatures quoted here were checked against that release. Blocking unsupported answers at runtime is a separate job, covered in ADK Java RAG Grounding, in depth; this article is about diagnosing why they happen.

A taxonomy of agent hallucinations

Most agent hallucinations are not the model inventing facts out of nowhere. They are the model doing something reasonable with bad or missing inputs. Debugging starts by naming the classes, because each one is fixed in a different place.

ClassWhat the trace showsLayer to fix
Skipped toolNo tool call, yet the answer states a fact the tool ownsInstruction, tool description
Fabricated argumentsTool called with an id that appears nowhere in the conversationTool schema, argument validation
Gap fillingTool returned empty, an error or NOT_FOUND; the answer states a value anywayTool result contract
Misread resultThe value is in the result but the answer contradicts itResult shape, field naming
Context lossThe tool result exists in the trace but not in the next model requestHistory, include-contents, compression
Prompt leakThe answer repeats an example value from the instructionInstruction examples
Parametric recallNothing in the turn mentions the claim; it came from training dataGrounding contract, verifier

In tool-using agents, parametric recall, the failure everyone pictures, is often not the main source, and it is the cheapest to block. Gap filling and skipped tools are the expensive ones, because the answer looks exactly like a correct answer.

Step 1: capture what the model actually saw

You cannot debug a hallucination from the final answer alone. You need four records per model step, keyed by the invocation id: the request the model actually received, its response, every tool call's arguments, and every tool result. ADK Java exposes all four as callbacks on LlmAgent. In 1.10.1 the synchronous signatures are: before-model receives CallbackContext and an LlmRequest.Builder; after-model receives CallbackContext and the LlmResponse; the tool callbacks receive an InvocationContext, the BaseTool, the argument map and a ToolContext, and after-tool also gets the result object.

public final class FlightRecorder {
  private final TraceSink sink;               // your log pipeline, table or file
  public FlightRecorder(TraceSink sink) { this.sink = sink; }

  public Optional<LlmResponse> beforeModel(CallbackContext ctx, LlmRequest.Builder req) {
    LlmRequest snapshot = req.build();        // read-only copy; the builder is not changed
    sink.write(ctx.invocationId(), "model_request", Map.of(
        "agent", ctx.agentName(),
        "tools", List.copyOf(snapshot.tools().keySet()),
        "contents", Json.of(snapshot.contents())));
    return Optional.empty();                  // never alter the request
  }

  public Optional<LlmResponse> afterModel(CallbackContext ctx, LlmResponse resp) {
    if (resp.partial().orElse(false)) return Optional.empty();
    sink.write(ctx.invocationId(), "model_response", Map.of(
        "finish", resp.finishReason().map(Object::toString).orElse("none"),
        "content", Json.of(resp.content())));
    return Optional.empty();
  }

  public Optional<Map<String, Object>> beforeTool(InvocationContext inv, BaseTool tool,
      Map<String, Object> args, ToolContext tc) {
    sink.write(inv.invocationId(), "tool_call", Map.of("tool", tool.name(), "args", Json.of(args)));
    return Optional.empty();                  // empty means: run the real tool
  }

  public Optional<Map<String, Object>> afterTool(InvocationContext inv, BaseTool tool,
      Map<String, Object> args, ToolContext tc, Object result) {
    sink.write(inv.invocationId(), "tool_result", Map.of("tool", tool.name(), "result", Json.of(result)));
    return Optional.empty();                  // empty means: keep the real result
  }
}

FlightRecorder rec = new FlightRecorder(sink);
LlmAgent agent = LlmAgent.builder()
    .name("order_support")
    .model("gemini-2.5-flash")
    .instruction(INSTRUCTION)
    .tools(FunctionTool.create(OrderTools.class, "getOrderStatus"))
    .beforeModelCallbackSync(rec::beforeModel)
    .afterModelCallbackSync(rec::afterModel)
    .beforeToolCallbackSync(rec::beforeTool)
    .afterToolCallbackSync(rec::afterTool)
    .build();

Three details matter. Snapshot the request in before-model, not the session history, because the request is what the model saw after ADK assembled instructions, history and tool declarations; differences between the two are exactly how context-loss bugs hide. Skip partial streaming responses so you record the final text once. And every callback returns Optional.empty(), so recording never changes behaviour. TraceSink and Json are your own helpers; redact personal data before the sink, because these traces contain everything the user and your systems said. The callback mechanics, including how several callbacks compose, are in ADK Java callback architecture, and the event model the session stores is in ADK Java Events, in depth.

Capture, replay, classify, fix: where each step hooks into an ADK Java agentUser turninvocationIdbeforeModelsnapshot requestGeminiLlmResponseafterModelsnapshot responseFunctionToolreal systembefore/afterToolargs and resultTrace storeper invocationReplay harnesstape, N runsClassifierroot-cause labelfunction callRecording never changes behaviour: every capture callback returns Optional.empty().
The debugging pipeline. Capture callbacks write per-invocation records; the replay harness feeds recorded tool results back; the classifier labels the root cause.

Step 2: replay until it reproduces

A hallucination you saw once is an anecdote. Before changing anything, reproduce it and measure how often it happens, because the fix has to be judged by a rate, not a single run. Build a tape from the trace: the user message, the instruction, the model id, and the recorded tool results. Then replay the turn with a before-tool callback that returns recorded results instead of calling live systems. In ADK, a before-tool callback that returns a non-empty map short-circuits the tool and that map becomes the result, which makes replay cheap and safe.

// Serve recorded tool results instead of calling live systems.
Callbacks.BeforeToolCallbackSync fromTape = (inv, tool, args, tc) ->
    Optional.of(tape.resultFor(tool.name(), args));   // non-empty result: the real tool is skipped

LlmAgent replay = LlmAgent.builder()
    .name("order_support")
    .model(tape.model())                               // the model id from the trace
    .instruction(tape.instruction())                   // the instruction text from the trace
    .tools(FunctionTool.create(OrderTools.class, "getOrderStatus"))
    .beforeToolCallbackSync(fromTape)
    .build();

InMemoryRunner runner = new InMemoryRunner(replay, "replay");
int reproduced = 0;
for (int i = 0; i < 20; i++) {
  Session s = runner.sessionService().createSession("replay", "debug").blockingGet();
  List<Event> events = runner.runAsync("debug", s.id(),
          Content.fromParts(Part.fromText(tape.userMessage())), RunConfig.builder().build())
      .toList().blockingGet();
  if (!Claims.unsupported(finalText(events), tape.evidence()).isEmpty()) reproduced++;
}
System.out.printf("reproduced %d of 20%n", reproduced);

Replay at production settings first to get the true rate; the incident happened at production temperature, and a temperature of 0.0f can hide it. When you later bisect a single change, lowering temperature through generateContentConfig reduces variance but does not guarantee identical output, so keep running N replays rather than one. If the failure never reproduces, the tape is incomplete: compare the replayed model request with the recorded one; missing earlier turns in the session are the usual culprit. Claims.unsupported is your claim checker; start with exact matching of numbers, ids and names against the evidence, then add an LLM judge as described in LLM-as-Judge Scorer in ADK Java.

Step 3: classify the root cause

With a reproducing tape, classify each unsupported claim with a fixed procedure. Doing it the same way every time lets you count classes across incidents, which tells you where to invest.

def classify(claim, trace):
    """claim: a statement in the final answer that no evidence in the turn supports."""
    owner = tool_that_owns(claim)                   # e.g. order status -> getOrderStatus
    if owner and owner not in trace.tools_called():
        return "SKIPPED_TOOL"
    for call in trace.calls(owner):
        if not trace.reached_model(call.result):
            return "CONTEXT_LOSS"                   # returned, but absent from the next request
        if any(v not in trace.user_text() + trace.prior_results() for v in call.id_args()):
            return "FABRICATED_ARGS"
        if call.result_is_empty_or_error():
            return "GAP_FILLING"
        if call.result_mentions(claim.subject):
            return "MISREAD"                        # the right field was there; answer disagrees
    if appears_in(claim, trace.instruction_examples()):
        return "PROMPT_LEAK"
    return "PARAMETRIC"

The order of checks matters. Context loss comes first among called tools because a result that never reached the model can be neither misread nor gap-filled. Fabricated arguments come before gap filling because an invented id usually produces an empty result, and fixing the empty-result handling would treat the symptom. reached_model compares the tool result with the recorded model request, which is why step 1 snapshots the request itself.

Worked example: the order that shipped but did not exist

A user writes: where is order 1043? The agent replies that the order shipped yesterday by courier with a tracking number, and none of that is true. The trace shows the model called getOrderStatus with orderId: "1043". The order system stores ids as A-1043, so the tool returned {"status": "NOT_FOUND"}. The next model request contains that result, and the response invents a shipped status.

Walk the procedure. The owning tool was called, its result reached the model, and the argument came from the user's text, so this is neither a skipped tool, context loss nor a fabricated argument. The result was a not-found, so the class is gap filling. The result shape explains why: a field called status holding NOT_FOUND reads like one order status among many rather than a terminal fact. Replaying the tape 20 times gives a baseline rate to beat.

The fix lands in two layers. First, the tool normalises the id, so 1043 finds A-1043. Second, the not-found shape becomes explicit and instructive:

return Map.of(
    "found", false,
    "orderId", orderId,
    "message", "No order matches this id. Do not guess a status. "
             + "Ask the user to confirm the full order id, for example A-1043.");

Rerun the 20 replays on a tape whose recorded result now uses the new shape; the fix counts when the reproduction rate drops to zero and a valid id still gets a correct answer. Then add the tape to your regression suite so a later prompt edit cannot quietly reintroduce the bug.

Fixes by layer

Each class has a fix at its own layer. Prompt edits are rarely the right first move.

  • Skipped tool. Make the tool description say which questions it owns, and state in the instruction that facts of that kind may only come from that tool. Check that the tool is actually in the request's tool list; a tool attached to the wrong sub-agent looks the same.
  • Fabricated arguments. Validate ids inside the tool and return a structured error naming the expected format. Never let a lookup silently succeed on a fuzzy match.
  • Gap filling. Return explicit found: false results with an instruction in the message, as above. Avoid null fields and empty lists without explanation.
  • Misread result. Flatten deep JSON, use unambiguous field names such as estimatedDeliveryDate, and drop fields the model does not need.
  • Context loss. Check history compression, includeContents(IncludeContents.NONE) on sub-agents, and truncation of large tool results; summarise big results inside the tool.
  • Prompt leak. Replace realistic example values in instructions with obvious placeholders, so a leaked example is detectable.
  • Parametric recall. Add a citation contract and a runtime verifier; that is enforcement, not debugging.

Operating it in production

In production, turn the workflow into metrics. Run the claim checker on a sample of turns and record the classifier label, so the dashboard shows rates by class, by tool and by agent rather than a single hallucination number. Alert on the gap-filling rate after empty tool results, which is the most damaging and the easiest to detect: the evidence is literally an empty result followed by a confident value. Keep every confirmed incident's tape in an evaluation set, as described in ADK Java Evaluation Framework, in depth, and rerun the set on every prompt, tool or model change.

Failure modes of the debugging process

  • Debugging from the session instead of the request. Session events and the model request differ after compression or include-contents settings; only the request snapshot is truth.
  • One-shot replay. A single clean replay proves nothing for a 1-in-10 failure. Use N runs.
  • Live tools during replay. Calling real systems changes the evidence and can repeat side effects such as refunds. Always replay from the tape.
  • Model drift. Record the model version; replaying against a newer model measures the model change, not your fix.
  • Traces as a data leak. Full request snapshots contain personal data; redact and set retention before turning recording on.

Trade-offs

ChoiceBenefitCost
Record every turnAny incident is debuggableStorage, privacy exposure
Record a sample plus all flagged turnsCheap, still covers incidentsUnflagged failures lack traces
Exact-match claim checkerDeterministic, fastMisses paraphrase
Judge-model claim checkerHandles paraphrase and inferenceCost, its own error rate
Fix in the toolFixes every caller and promptRequires code change and deploy
Fix in the promptFast to shipFragile across model versions

What to do next

  1. Add the FlightRecorder callbacks to one agent, with redaction, and confirm one trace contains request, response, tool calls and results.
  2. Write a tape builder from a trace and a replay harness that serves tool results through a before-tool callback.
  3. Implement exact-match Claims.unsupported for numbers, ids, dates and names.
  4. Take your last three hallucination reports, replay each 20 times, and label them with the classifier.
  5. Fix each at its layer, rerun the replays, and add the tapes to your evaluation set.
  6. Change every tool's empty and not-found results to explicit, instructive shapes, then add a gap-filling-rate alert.
Key takeaway: Treat each hallucination as a bug with a layer. Record the exact model request, response, tool arguments and results per invocation; replay from recorded tool results until you have a reproduction rate; classify the claim as skipped tool, fabricated arguments, gap filling, misread, context loss, prompt leak or parametric recall; and fix it in the tool, the result shape or the context before reaching for the prompt.