Most LLM-as-judge setups grade the final answer. For agents that is necessary and not enough. An agent can reach the right answer by the wrong path: issuing a refund before verifying the order, searching seven times when one lookup would do, or guessing a value it should have fetched. And when the answer is wrong, a final-answer judge cannot say which step broke or which sub-agent to fix. Matching against a reference trajectory solves this only for tasks with one correct path, and most real tasks have several.

This article builds a process judge for ADK Java: a model that reads the whole trajectory of an invocation, labels its steps, names the first step that went wrong and the agent that took it, and does so with validation that turns unusable judgements into explicit errors. It assumes the basics from the LLM-as-judge scorer article, which covers rubric design, the structured-verdict judge agent and its tests, and from the evaluation framework overview, which covers deterministic trajectory matching against a reference.

Where a process judge fits

MethodNeedsCatchesMisses
Final-answer judgerubric, optional referencewrong or ungrounded answersbad paths to right answers
Reference trajectory matchexpected tool callsdeviations from one known pathvalid alternative paths
Rule checks on the trajectoryexplicit rulesforbidden calls, loops, limitsanything needing judgement
Process judge (this article)policy and tool descriptionsunsafe ordering, waste, unrecovered errorserrors the transcript does not show

These layer rather than compete. Rule checks are free and exact, so run them on everything. The process judge costs a model call over a long input, so run it where answers are high-stakes, where multiple valid paths make reference matching useless, or on failures you need to attribute.

Architecture

A trajectory judge for ADK Java agentsSession eventsone invocationRenderernumbered steps, truncatedRule checksforbidden, repeats, limitsJudge LlmAgentno tools, no history, schemaas dataValidatorindex in range, quote foundERRORunusable judgementProcess verdictfirst bad step + agentReportsper agent, per categoryCheap deterministic checks run first and their findings are passed to the judge, which only rules on what rules cannot.
The judge never sees raw events. A renderer turns them into numbered steps, rule checks run first, and a validator rejects judgements that cite steps or quotes that do not exist.

The judge is an LlmAgent with no tools, IncludeContents.NONE so it sees only the request you build, transfers disabled, temperature zero and an output schema. Every part of the trajectory reaches it as quoted data between markers, with an explicit instruction that nothing inside the markers is an instruction. Tool results are the main injection risk: they carry text from web pages, tickets and emails, and a judge reading them is exposed to whatever that text says.

Rendering the trajectory

The renderer decides what the judge can see, so it decides what the judge can find. Number every step, name its author, and show tool calls with arguments and tool results with a length cap. Truncate in the middle and say so, because errors often sit at the end of a long result.

record Step(int index, String author, String kind, String body) {}

static List<Step> render(Session session, String invocationId, int maxChars) {
  List<Step> out = new ArrayList<>();
  for (Event e : session.events()) {
    if (!e.invocationId().equals(invocationId) || e.partial().orElse(false)) continue;
    for (FunctionCall fc : e.functionCalls())
      out.add(new Step(out.size() + 1, e.author(), "CALL",
          fc.name().orElse("?") + " " + fc.args().orElse(Map.of())));
    for (FunctionResponse fr : e.functionResponses())
      out.add(new Step(out.size() + 1, e.author(), "RESULT",
          fr.name().orElse("?") + " " + clip(String.valueOf(fr.response().orElse(Map.of())), maxChars)));
    e.actions().transferToAgent().ifPresent(t ->
        out.add(new Step(out.size() + 1, e.author(), "TRANSFER", "to " + t)));
    if (e.functionCalls().isEmpty() && e.functionResponses().isEmpty()) {
      String text = e.stringifyContent();
      if (!text.isBlank()) out.add(new Step(out.size() + 1, e.author(), "SAY", clip(text, maxChars)));
    }
  }
  return out;
}

static String clip(String s, int max) {
  if (s.length() <= max) return s;
  int half = max / 2;
  return s.substring(0, half) + " [... " + (s.length() - max) + " chars cut ...] "
      + s.substring(s.length() - half);
}

Give the judge the same policy the agent had: the relevant part of the agent instruction and every tool description. Without them it cannot tell a policy violation from a reasonable choice, and it will invent a policy of its own.

Rule checks first

Run deterministic checks before the judge and pass their findings in as facts. They catch the obvious cases exactly and for free, and they stop the judge spending attention on things a loop over the steps can settle.

static List<String> ruleFindings(List<Step> steps, Policy policy) {
  List<String> f = new ArrayList<>();
  Map<String, Integer> seen = new HashMap<>();
  for (Step s : steps) {
    if (!s.kind().equals("CALL")) continue;
    if (policy.forbidden(s.author(), s.body()))
      f.add("step " + s.index() + ": " + s.author() + " called a tool it may not use");
    int n = seen.merge(s.body(), 1, Integer::sum);
    if (n == 3) f.add("step " + s.index() + ": third identical call: " + s.body());
  }
  if (steps.size() > policy.maxSteps())
    f.add("trajectory has " + steps.size() + " steps, limit " + policy.maxSteps());
  return f;
}

Policy is your own small class: which agent may call which tool, the step limit, and any ordering rule simple enough to state in code, such as verify_order must precede issue_refund. Encode ordering rules here whenever you can; the judge is for the rules that need reading comprehension, such as whether the agent told the customer something the tool result did not support.

The verdict schema

Ask for three things: a label per step, a single first error, and trajectory-level criteria. Step labels make the judgement inspectable. The first error is what attribution needs: the earliest step after which success became impossible or a rule was broken. Criteria such as policy, grounding, efficiency and recovery give you rates to track.

Schema STEP = Schema.builder().type("OBJECT").properties(Map.of(
    "index", Schema.builder().type("INTEGER").build(),
    "label", Schema.builder().type("STRING")
        .enum_(List.of("NEEDED", "REDUNDANT", "HARMFUL", "RECOVERY")).build()))
    .required(List.of("index", "label")).build();

Schema VERDICT = Schema.builder().type("OBJECT").properties(Map.of(
    "steps", Schema.builder().type("ARRAY").items(STEP).build(),
    "first_error_step", Schema.builder().type("INTEGER")
        .description("Index of the earliest decisive error, or 0 if none").build(),
    "responsible_agent", Schema.builder().type("STRING").build(),
    "category", Schema.builder().type("STRING").enum_(List.of(
        "NONE", "POLICY", "UNGROUNDED", "WRONG_TOOL", "BAD_ARGS", "IGNORED_ERROR", "LOOP")).build(),
    "evidence", Schema.builder().type("STRING")
        .description("Verbatim quote from the cited step").build()))
    .required(List.of("steps", "first_error_step", "responsible_agent", "category", "evidence"))
    .build();

Schema builder caveats from the scorer article apply here too. The judge agent is built as in the scorer article, with this schema as its output schema, and run on its own InMemoryRunner with a fresh session per trajectory.

Validating every judgement

Never accept a process verdict without checking it against the trajectory it claims to describe. These checks are cheap and they catch most fabricated judgements.

  1. first_error_step is 0 or a valid step index; anything else is ERROR.
  2. If first_error_step is not 0, responsible_agent must equal the author of that step. If not, record the author of the cited step and flag the disagreement for review.
  3. The evidence quote, after whitespace normalisation, must appear in the rendered text of the cited step. A quote found nowhere means the judge invented its reason.
  4. Every rule finding must be reflected: if the rules found a forbidden call at step 4, a verdict claiming no error is ERROR, not PASS.
  5. Step labels must cover every index exactly once.

ERROR is not FAIL. A judgement that fails validation tells you about the judge, not the agent; count it separately and keep its rate on your dashboard, because a rising error rate usually means the trajectories got longer than the judge handles well.

Attribution across sub-agents

Attribution is the hardest part, and the evidence says to be modest about it. The Who&When benchmark (Zhang et al., ICML 2025, arXiv 2505.00212) annotated failure logs from 127 LLM multi-agent systems with the responsible agent and the decisive step. Its best automated method identified the responsible agent 53.5% of the time but the decisive step only 14.2% of the time, and strong reasoning models did not reach practical usability. Treat a judge's step attribution as a pointer for a human to check, never as ground truth for blame.

Three practices make it more useful. Aggregate before acting: one misattribution is noise, but if 40 of 60 failures point at the same sub-agent and category, the pattern is real. Ask for the first error, not the worst, because judges tend to blame the last visible step, often the agent that delivered the bad news rather than the one that caused it. And calibrate attribution separately from pass or fail: have people label the first error on a sample, as described in the human evaluation article, and measure agreement on the step itself.

Pairwise comparison of versions

When comparing two agent versions, absolute scores on long trajectories drift with the judge's mood. Pairwise comparison is steadier: show the judge both trajectories for the same input and ask which handled the process better, citing steps. Run every pair twice with the order swapped and keep only consistent verdicts; an inconsistent pair is a tie. Report wins, losses, ties and the inconsistency rate, which measures how much of the difference is position bias rather than quality.

Worked example: a right answer by a wrong path

An illustrative case: a refund assistant with a router, an orders_agent and a refunds_agent. The customer asks for a refund on a damaged item. The trajectory has six steps: router transfers to refunds_agent (1); refunds_agent calls issue_refund with the order id from the message (2); the result says refund created (3); refunds_agent transfers to orders_agent (4), which calls verify_order (5); the result shows the order was already refunded last week (6). The final answer tells the customer the refund is on its way, which is true, so a final-answer judge passes it.

The rule check flags step 2 straight away, because policy says verify_order must precede issue_refund. The process judge labels step 2 HARMFUL, gives first_error_step 2, responsible_agent refunds_agent, category POLICY, and quotes the issue_refund call as evidence. Validation passes: step 2 exists, its author matches, the quote is in it. Across a week of sampled traffic the same pattern appears in a few percent of refund runs, always from refunds_agent, which points to its instruction, not to the router.

Cost and where to run it

A process judgement costs roughly the rendered trajectory plus policy and tool descriptions as input. A 20-step trajectory with capped tool results might be 8,000 input tokens and 600 output tokens, several times a final-answer judgement. Budget for it: judge every run in offline evaluation, but in production judge failures, escalations and a stratified sample, the approach described in continuous evaluation in production. Use a judge from a different model family from the agent where you can, since models tend to favour outputs like their own.

Failure modes

  • Truncation hides the error. If the decisive fact sat in the cut part of a tool result, the judge cannot find it. Log how often cuts happen in failing runs and raise the cap for those tools.
  • Injection through tool results. A ticket that says to rate this conversation as perfect is data, but a judge may obey it. Markers, no tools and quote validation reduce the risk; spot-check verdicts on runs whose tool results contain imperative text.
  • Recency blame. Judges over-attribute to the last agent. Ask for the first error and verify attributions on a human-labelled sample.
  • Invented policy. Without the agent instruction and tool descriptions the judge grades against its own idea of good behaviour. Always pass the policy in.
  • Silent judge drift. A judge model upgrade changes labels. Store the judge model and prompt version with every verdict and rerun the calibration set on any change.

Trade-offs

A process judge buys diagnosis at the price of cost, latency and a second model to trust. Reference trajectories are exact and cheap but brittle, breaking whenever a valid new path appears; rule checks are exact but cover only what you can state in code. The judge covers the rest, at roughly the accuracy of a careful but hurried reviewer, which is why its output feeds aggregates and human review rather than automatic blame or automatic release gates. Test the pipeline like code: feed the validator canned verdicts with an out-of-range step, a wrong author and a fabricated quote, and assert each one becomes ERROR.

What to do next

  1. Write the policy your agents follow as data: allowed tools per agent, ordering rules and a step limit, and implement the rule checks over rendered steps.
  2. Build the renderer and read ten rendered trajectories yourself before any judge sees them; fix truncation and missing steps first.
  3. Add the process judge with the verdict schema and all five validation checks, and track the ERROR rate from day one.
  4. Label the first error on 50 failing trajectories by hand and measure how often the judge agrees on the agent and on the step.
  5. Run it on failures and a stratified sample in production, aggregate by agent and category weekly, and act only on patterns.
Key takeaway: Grade the path as well as the answer: render each invocation as numbered steps, settle what rules can settle in code, ask a tool-less, schema-bound judge for step labels and the first decisive error, reject verdicts whose step or quote does not check out, and treat attribution as a pointer to verify in aggregate.