An agent evaluation framework answers one question before every release: does this version still do the right things on the conversations we care about? For an ADK Java agent that means two separate checks: which tools it called, with which arguments and in what order, and what it finally said. This page builds the framework: a portable eval-set format, a replay harness, deterministic scorers and a CI gate.

You have to build it yourself today. The Python ADK ships an evaluator, an adk eval command and a set of named criteria. ADK Java, as of v1.11.0 (released 2 October 2026), has no evaluator in its core module, and the dev server's EvaluationController routes for eval sets and run-eval are placeholders that log a warning and return NOT_IMPLEMENTED or empty lists. The request DTOs are there (RunEvalRequest carries evalIds and evalMetrics), so check each new release.

The framework in one picture

*.evalset.jsoncases, turns, expected toolsLoadersnake_case or camelCasereadHarnessfresh InMemoryRunner per caseEvalCaseAgent under testreal or scripted BaseLlmrunAsyncEvent streamfunctionCalls, finalResponseeventsTrajectory scorerEXACT, IN_ORDER, ANY_ORDERResponse scorerunigram F1Judge scorerrubric, separate pageBudget checkslatency, tool count, errorsAggregator and gateper-case thresholds, repeats, JSON report, JUnit resultRecorded turns go in; one verdict per case and one release decision come out.
Eval cases are replayed through the real agent on a fresh runner; the event stream feeds independent scorers, and the aggregator turns their scores into a gate and a report.

What exists today and what you build

An eval set is a file of eval cases. A case is one conversation: one or more turns, each with the user's message, the expected final response and the expected tool calls, plus the initial session state. An invocation is what the agent did in one turn, and a criterion scores invocations against a threshold.

PiecePython ADKADK Java v1.11.0What this page builds
Eval-set file formatEvalSet and EvalCase models, *.evalset.jsonNoneA loader that reads the Python format
Runner for casesadk eval command and pytest helperDev-server routes are placeholdersA harness on InMemoryRunner
Tool trajectory criteriontool_trajectory_avg_score, three match typesNoneThe same three match types
Response criterionresponse_match_score, ROUGE-1NoneUnigram F1, documented as not identical
Model-judged criteriaSeveral, including rubric-based onesNoneLinked: a rubric judge on LlmAgent

Reusing the Python format means teams running both languages share one set of cases, and when ADK Java ships an evaluator your files are the most likely ones to load.

A portable eval-set format

The Python ADK's eval-set file has a small, stable core: eval_set_id and eval_cases at the top; per case an eval_id, a conversation list and an optional session_input; per turn an invocation_id, user_content, final_response and intermediate_data holding tool_uses with a name and arguments. The samples in the adk-python repository use snake_case keys; files saved by other tooling may use camelCase aliases of the same names, so the loader accepts either.

{
  "eval_set_id": "home_automation",
  "eval_cases": [{
    "eval_id": "list_then_turn_off",
    "conversation": [{
      "invocation_id": "t1",
      "user_content":   {"role": "user",  "parts": [{"text": "Which devices are on?"}]},
      "final_response": {"role": "model", "parts": [{"text": "The Living Room device (device_1) is on."}]},
      "intermediate_data": {"tool_uses": [{"name": "list_devices", "args": {"status": "ON"}}]}
    }],
    "session_input": {"app_name": "home_automation_agent", "user_id": "user", "state": {}}
  }]
}

Read it into plain records holding only what the scorers use; unknown fields are ignored, so files from newer tools still load.

public record ToolUse(String name, Map<String, Object> args) {}
public record Turn(String id, String userText, String expectedText, List<ToolUse> expectedTools) {}
public record EvalCase(String id, List<Turn> turns, Map<String, Object> state) {}

final class EvalSetLoader {
  private static final ObjectMapper JSON = new ObjectMapper();
  private static final TypeReference<Map<String, Object>> MAP = new TypeReference<>() {};

  /** Accept both spellings: files written by different tools disagree. */
  private static JsonNode f(JsonNode n, String snake, String camel) {
    JsonNode v = n.get(snake);
    return v != null ? v : n.path(camel);
  }

  private static String text(JsonNode content) {
    StringBuilder sb = new StringBuilder();
    for (JsonNode part : content.path("parts")) sb.append(part.path("text").asText(""));
    return sb.toString();
  }

  private static Map<String, Object> obj(JsonNode n) {
    return n.isObject() ? Normalize.map(JSON.convertValue(n, MAP)) : Map.of();
  }

  static List<EvalCase> load(Path file) throws IOException {
    List<EvalCase> cases = new ArrayList<>();
    for (JsonNode c : f(JSON.readTree(file.toFile()), "eval_cases", "evalCases")) {
      List<Turn> turns = new ArrayList<>();
      for (JsonNode t : c.path("conversation")) {
        List<ToolUse> tools = new ArrayList<>();
        for (JsonNode u : f(f(t, "intermediate_data", "intermediateData"), "tool_uses", "toolUses")) {
          tools.add(new ToolUse(u.path("name").asText(), obj(u.path("args"))));
        }
        turns.add(new Turn(f(t, "invocation_id", "invocationId").asText(""),
            text(f(t, "user_content", "userContent")),
            text(f(t, "final_response", "finalResponse")), tools));
      }
      cases.add(new EvalCase(f(c, "eval_id", "evalId").asText(), turns,
          obj(f(c, "session_input", "sessionInput").path("state"))));
    }
    return cases;
  }
}

Replaying a case through the agent

The harness replays each case through the real agent, with its real instructions, tools and callbacks. Give every case a fresh InMemoryRunner and session so nothing leaks between cases, seed the session with the case's state, and feed turns in order through that session so turn two sees turn one's history.

Everything the scorers need is in the event stream. Runner.runAsync(userId, sessionId, content) returns a Flowable<Event>; each event's functionCalls() lists the tool calls the model requested, and finalResponse() is true for the events that end the turn. Collect calls from every event and text only from final ones. Sub-agents' calls arrive in the same stream with a different author(); decide whether expectations include them.

public record Invocation(String finalText, List<ToolUse> tools, long latencyMs, String error) {}

static List<Invocation> replay(BaseAgent agent, EvalCase c) {
  InMemoryRunner runner = new InMemoryRunner(agent);          // fresh services per case
  String user = "eval-user";
  Session session = runner.sessionService()
      .createSession(runner.appName(), user, new HashMap<>(c.state()), null)
      .blockingGet();
  List<Invocation> out = new ArrayList<>();
  for (Turn t : c.turns()) {
    long start = System.nanoTime();
    List<ToolUse> tools = new ArrayList<>();
    StringBuilder text = new StringBuilder();
    String error = null;
    try {
      Content msg = Content.fromParts(Part.fromText(t.userText()));
      for (Event e : runner.runAsync(user, session.id(), msg).blockingIterable()) {
        for (FunctionCall fc : e.functionCalls()) {
          tools.add(new ToolUse(fc.name().orElse(""), Normalize.map(fc.args().orElse(Map.of()))));
        }
        if (e.finalResponse()) {
          e.content().flatMap(Content::parts)
              .ifPresent(ps -> ps.forEach(p -> p.text().ifPresent(text::append)));
        }
      }
    } catch (RuntimeException ex) {
      error = ex.toString();                                  // an ERROR, not a FAIL
    }
    out.add(new Invocation(text.toString(), tools, (System.nanoTime() - start) / 1_000_000, error));
  }
  return out;
}

The catch block records exceptions as errors, reported separately from wrong answers: an outage should page someone, not look like a quality regression.

Scoring the tool trajectory

The Python ADK's tool_trajectory_avg_score defines three match types worth copying exactly. EXACT demands the same calls in the same order with nothing extra. IN_ORDER requires the expected calls in order, with other calls allowed between them. ANY_ORDER requires each expected call somewhere. Each turn scores 1 or 0, and the case score is the mean over its turns, compared against a threshold that defaults to 1.0 in the Python ADK.

Argument comparison breaks naive implementations: the same value can arrive as Integer 3 on one side and Double 3.0 on the other. Normalise both sides before comparing.

enum MatchType { EXACT, IN_ORDER, ANY_ORDER }

static boolean trajectoryMatches(List<ToolUse> expected, List<ToolUse> actual, MatchType type) {
  switch (type) {
    case EXACT:
      return expected.equals(actual);
    case IN_ORDER: {                       // expected is a subsequence of actual
      int i = 0;
      for (ToolUse a : actual) if (i < expected.size() && expected.get(i).equals(a)) i++;
      return i == expected.size();
    }
    default: {                             // ANY_ORDER: multiset containment
      List<ToolUse> pool = new ArrayList<>(actual);
      for (ToolUse e : expected) if (!pool.remove(e)) return false;
      return true;
    }
  }
}

final class Normalize {
  /** Whole numbers become Long, others Double, maps sorted, so 3 equals 3.0. */
  static Object value(Object v) {
    if (v instanceof Number n) {
      double d = n.doubleValue();
      return (d == Math.rint(d) && !Double.isInfinite(d)) ? (Object) (long) d : (Object) d;
    }
    if (v instanceof Map<?, ?> m) return map(m);
    if (v instanceof List<?> l) return l.stream().map(Normalize::value).toList();
    return v;
  }
  static Map<String, Object> map(Map<?, ?> m) {
    Map<String, Object> r = new TreeMap<>();
    m.forEach((k, x) -> r.put(String.valueOf(k), value(x)));
    return r;
  }
}

Scoring the final response

The Python ADK's response_match_score is documented as ROUGE-1 similarity, meaning unigram overlap between the response and the reference, with a default threshold of 0.8. The Java version below computes unigram F1 with clipped counts. Tokenisation and stemming differ between implementations, so do not assume it equals the Python score, and set your own threshold.

static double unigramF1(String reference, String candidate) {
  Map<String, Integer> ref = counts(reference), cand = counts(candidate);
  int overlap = 0;
  for (var e : cand.entrySet()) overlap += Math.min(e.getValue(), ref.getOrDefault(e.getKey(), 0));
  if (overlap == 0) return 0.0;
  double precision = (double) overlap / total(cand), recall = (double) overlap / total(ref);
  return 2 * precision * recall / (precision + recall);
}

static Map<String, Integer> counts(String s) {
  Map<String, Integer> m = new HashMap<>();
  for (String tok : s.toLowerCase(Locale.ROOT).split("[^a-z0-9]+")) {
    if (!tok.isEmpty()) m.merge(tok, 1, Integer::sum);
  }
  return m;
}

static int total(Map<String, Integer> m) {
  return m.values().stream().mapToInt(Integer::intValue).sum();
}

Worked example: one case, three scores

Take the case above with one more turn: the user says "Turn that one off." and the expected call is set_device_info with device_id device_1 and status OFF. The agent under test first calls get_user_prefs with no arguments, then list_devices with status ON, and replies "Device device_1 in the Living Room is currently on." On turn two it calls set_device_info with the same two arguments in the opposite key order.

CheckTurn 1Turn 2Case score
Trajectory, EXACT0 (extra call)1 (key order is irrelevant)0.5, fails at 1.0
Trajectory, IN_ORDER111.0, passes
Response F10.889scored separatelypasses at 0.8

By hand: the reference tokenises to eight words: the, living, room, device, device, 1, is, on. The response gives ten, and all eight reference words are among them, so precision is 8/10 = 0.8, recall is 8/8 = 1.0 and F1 is 2 x 0.8 x 1.0 / 1.8 = 0.889.

Now the warning. A response of "The Living Room device (device_1) is off." shares seven of its eight words with the reference, so precision and recall are both 0.875 and it also passes at 0.8, while being factually wrong. Lexical overlap measures phrasing, not truth. Treat it as a tripwire, rely on the trajectory for correct lookups, and add a rubric judge for meaning; this site's LLM-as-judge scorer for ADK Java builds one on LlmAgent with a structured verdict.

Determinism: scripted models and repeats

Test the harness and scorers deterministically: subclass BaseLlm, whose abstract methods are generateContent(LlmRequest, boolean) returning Flowable<LlmResponse> and connect(LlmRequest), and feed it canned responses. The custom LLM guide covers that class in detail.

final class ScriptedLlm extends BaseLlm {
  private final Deque<LlmResponse> script;
  ScriptedLlm(List<LlmResponse> responses) {
    super("scripted");
    this.script = new ArrayDeque<>(responses);
  }
  @Override public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
    LlmResponse next = script.poll();
    return next == null ? Flowable.error(new IllegalStateException("script exhausted"))
                        : Flowable.just(next);
  }
  @Override public BaseLlmConnection connect(LlmRequest request) {
    throw new UnsupportedOperationException("live mode not scripted");
  }
}

With a real model, even at temperature zero, one run is a sample, not a measurement. Run each case three to five times and decide in advance whether every repeat must pass or a majority suffices. Keep every repeat's trajectory: a case passing 4 of 5 times usually has one alternative path worth studying.

Wiring it into JUnit and CI

A JUnit 5 dynamic test per case gives per-case reporting for free. Keep a scripted suite for every commit and a live suite, run on merges or nightly, that holds the release gate.

class HomeAgentEvalTest {
  static final int REPEATS = 3;
  static final double F1_MIN = 0.6;          // set from your own baseline, not copied from Python

  @TestFactory
  Stream<DynamicTest> evalSet() throws IOException {
    return EvalSetLoader.load(Path.of("src/test/resources/evals/home.evalset.json")).stream()
        .map(c -> DynamicTest.dynamicTest(c.id(), () -> {
          for (int r = 0; r < REPEATS; r++) {
            List<Invocation> got = Harness.replay(HomeAgent.build(), c);
            for (int i = 0; i < got.size(); i++) {
              Invocation inv = got.get(i);
              Turn want = c.turns().get(i);
              assertNull(inv.error(), () -> c.id() + " errored: " + inv.error());
              assertTrue(Scorers.trajectoryMatches(want.expectedTools(), inv.tools(), MatchType.IN_ORDER),
                  () -> c.id() + " turn " + want.id() + " tools " + inv.tools());
              assertTrue(Scorers.unigramF1(want.expectedText(), inv.finalText()) >= F1_MIN,
                  () -> c.id() + " turn " + want.id() + " said: " + inv.finalText());
            }
          }
        }));
  }
}

Also write a JSON report with each case and repeat's tool calls, text, scores, latency, model and agent version; it is what you diff when the gate fails. Treat latency and tool-call counts as budgets: five calls where there used to be two is a regression even when the answer is right. The production counterpart of those numbers is covered in tool observability and metrics, and guardrail logic you may want to exercise from eval cases lives in ADK Java callbacks.

Failure modes

  • State leaking between cases. Results depend on case order. Use one runner per case.
  • Number types in arguments. Integer against Double breaks EXACT matching on every numeric argument. Normalise both sides.
  • Errors counted as failures, or as passes. A quota error is neither. Report errors separately and fail the gate on the error rate.
  • Over-specified trajectories. EXACT on an agent that looks things up fails every release. Prefer IN_ORDER.
  • Trusting lexical scores. F1 passes wrong answers that reuse the reference's words, as the worked example shows. Pair it with trajectory checks and a judge.
  • Expected outputs that rot. Changing tool data fails cases for the wrong reason. Use fixed tool fixtures.
  • One run per case. A single pass hides a flaky path.

Trade-offs

ChoiceGainsCosts
Python eval-set formatShared cases, likely forward compatibilityNested JSON, fields you do not use yet
Scripted modelFast, free, deterministicTests the harness, not the agent's judgment
Live model in the gateMeasures real behaviourCost, latency, flakiness, provider drift
Strict match typesCatch procedure changesFalse failures on harmless variation

What to do next

  1. Write ten eval cases from real conversations in the Python eval-set format.
  2. Implement the loader, harness and scorers, and unit-test them with the scripted model, including numeric arguments.
  3. Choose a match type per case, defaulting to IN_ORDER.
  4. Run the live suite five times per case to measure your baseline, then set the F1 threshold and the repeat rule from that data.
  5. Add a rubric judge for the criteria that matter most, and report errors separately from failures.
  6. Run the live suite in CI on merges, publish the JSON report, and check new ADK Java releases for native evaluation.
Key takeaway: ADK Java has no evaluator yet, but every piece one needs is a documented primitive. Load cases in the Python eval-set format, replay each on a fresh InMemoryRunner, score tool trajectories with explicit match types and normalised arguments, treat lexical response scores as a tripwire rather than proof, test the harness with a scripted model, repeat live runs, and gate releases on per-case results with errors reported separately.