An agent evaluation framework answers one question before every release: does this version still do the right things on the conversations we care about? For an ADK Java agent that means two separate checks: which tools it called, with which arguments and in what order, and what it finally said. This page builds the framework: a portable eval-set format, a replay harness, deterministic scorers and a CI gate.
You have to build it yourself today. The Python ADK ships an evaluator, an adk eval command and a set of named criteria. ADK Java, as of v1.11.0 (released 2 October 2026), has no evaluator in its core module, and the dev server's EvaluationController routes for eval sets and run-eval are placeholders that log a warning and return NOT_IMPLEMENTED or empty lists. The request DTOs are there (RunEvalRequest carries evalIds and evalMetrics), so check each new release.
The framework in one picture
What exists today and what you build
An eval set is a file of eval cases. A case is one conversation: one or more turns, each with the user's message, the expected final response and the expected tool calls, plus the initial session state. An invocation is what the agent did in one turn, and a criterion scores invocations against a threshold.
| Piece | Python ADK | ADK Java v1.11.0 | What this page builds |
|---|---|---|---|
| Eval-set file format | EvalSet and EvalCase models, *.evalset.json | None | A loader that reads the Python format |
| Runner for cases | adk eval command and pytest helper | Dev-server routes are placeholders | A harness on InMemoryRunner |
| Tool trajectory criterion | tool_trajectory_avg_score, three match types | None | The same three match types |
| Response criterion | response_match_score, ROUGE-1 | None | Unigram F1, documented as not identical |
| Model-judged criteria | Several, including rubric-based ones | None | Linked: a rubric judge on LlmAgent |
Reusing the Python format means teams running both languages share one set of cases, and when ADK Java ships an evaluator your files are the most likely ones to load.
A portable eval-set format
The Python ADK's eval-set file has a small, stable core: eval_set_id and eval_cases at the top; per case an eval_id, a conversation list and an optional session_input; per turn an invocation_id, user_content, final_response and intermediate_data holding tool_uses with a name and arguments. The samples in the adk-python repository use snake_case keys; files saved by other tooling may use camelCase aliases of the same names, so the loader accepts either.
{
"eval_set_id": "home_automation",
"eval_cases": [{
"eval_id": "list_then_turn_off",
"conversation": [{
"invocation_id": "t1",
"user_content": {"role": "user", "parts": [{"text": "Which devices are on?"}]},
"final_response": {"role": "model", "parts": [{"text": "The Living Room device (device_1) is on."}]},
"intermediate_data": {"tool_uses": [{"name": "list_devices", "args": {"status": "ON"}}]}
}],
"session_input": {"app_name": "home_automation_agent", "user_id": "user", "state": {}}
}]
}Read it into plain records holding only what the scorers use; unknown fields are ignored, so files from newer tools still load.
public record ToolUse(String name, Map<String, Object> args) {}
public record Turn(String id, String userText, String expectedText, List<ToolUse> expectedTools) {}
public record EvalCase(String id, List<Turn> turns, Map<String, Object> state) {}
final class EvalSetLoader {
private static final ObjectMapper JSON = new ObjectMapper();
private static final TypeReference<Map<String, Object>> MAP = new TypeReference<>() {};
/** Accept both spellings: files written by different tools disagree. */
private static JsonNode f(JsonNode n, String snake, String camel) {
JsonNode v = n.get(snake);
return v != null ? v : n.path(camel);
}
private static String text(JsonNode content) {
StringBuilder sb = new StringBuilder();
for (JsonNode part : content.path("parts")) sb.append(part.path("text").asText(""));
return sb.toString();
}
private static Map<String, Object> obj(JsonNode n) {
return n.isObject() ? Normalize.map(JSON.convertValue(n, MAP)) : Map.of();
}
static List<EvalCase> load(Path file) throws IOException {
List<EvalCase> cases = new ArrayList<>();
for (JsonNode c : f(JSON.readTree(file.toFile()), "eval_cases", "evalCases")) {
List<Turn> turns = new ArrayList<>();
for (JsonNode t : c.path("conversation")) {
List<ToolUse> tools = new ArrayList<>();
for (JsonNode u : f(f(t, "intermediate_data", "intermediateData"), "tool_uses", "toolUses")) {
tools.add(new ToolUse(u.path("name").asText(), obj(u.path("args"))));
}
turns.add(new Turn(f(t, "invocation_id", "invocationId").asText(""),
text(f(t, "user_content", "userContent")),
text(f(t, "final_response", "finalResponse")), tools));
}
cases.add(new EvalCase(f(c, "eval_id", "evalId").asText(), turns,
obj(f(c, "session_input", "sessionInput").path("state"))));
}
return cases;
}
}
Replaying a case through the agent
The harness replays each case through the real agent, with its real instructions, tools and callbacks. Give every case a fresh InMemoryRunner and session so nothing leaks between cases, seed the session with the case's state, and feed turns in order through that session so turn two sees turn one's history.
Everything the scorers need is in the event stream. Runner.runAsync(userId, sessionId, content) returns a Flowable<Event>; each event's functionCalls() lists the tool calls the model requested, and finalResponse() is true for the events that end the turn. Collect calls from every event and text only from final ones. Sub-agents' calls arrive in the same stream with a different author(); decide whether expectations include them.
public record Invocation(String finalText, List<ToolUse> tools, long latencyMs, String error) {}
static List<Invocation> replay(BaseAgent agent, EvalCase c) {
InMemoryRunner runner = new InMemoryRunner(agent); // fresh services per case
String user = "eval-user";
Session session = runner.sessionService()
.createSession(runner.appName(), user, new HashMap<>(c.state()), null)
.blockingGet();
List<Invocation> out = new ArrayList<>();
for (Turn t : c.turns()) {
long start = System.nanoTime();
List<ToolUse> tools = new ArrayList<>();
StringBuilder text = new StringBuilder();
String error = null;
try {
Content msg = Content.fromParts(Part.fromText(t.userText()));
for (Event e : runner.runAsync(user, session.id(), msg).blockingIterable()) {
for (FunctionCall fc : e.functionCalls()) {
tools.add(new ToolUse(fc.name().orElse(""), Normalize.map(fc.args().orElse(Map.of()))));
}
if (e.finalResponse()) {
e.content().flatMap(Content::parts)
.ifPresent(ps -> ps.forEach(p -> p.text().ifPresent(text::append)));
}
}
} catch (RuntimeException ex) {
error = ex.toString(); // an ERROR, not a FAIL
}
out.add(new Invocation(text.toString(), tools, (System.nanoTime() - start) / 1_000_000, error));
}
return out;
}The catch block records exceptions as errors, reported separately from wrong answers: an outage should page someone, not look like a quality regression.
Scoring the tool trajectory
The Python ADK's tool_trajectory_avg_score defines three match types worth copying exactly. EXACT demands the same calls in the same order with nothing extra. IN_ORDER requires the expected calls in order, with other calls allowed between them. ANY_ORDER requires each expected call somewhere. Each turn scores 1 or 0, and the case score is the mean over its turns, compared against a threshold that defaults to 1.0 in the Python ADK.
Argument comparison breaks naive implementations: the same value can arrive as Integer 3 on one side and Double 3.0 on the other. Normalise both sides before comparing.
enum MatchType { EXACT, IN_ORDER, ANY_ORDER }
static boolean trajectoryMatches(List<ToolUse> expected, List<ToolUse> actual, MatchType type) {
switch (type) {
case EXACT:
return expected.equals(actual);
case IN_ORDER: { // expected is a subsequence of actual
int i = 0;
for (ToolUse a : actual) if (i < expected.size() && expected.get(i).equals(a)) i++;
return i == expected.size();
}
default: { // ANY_ORDER: multiset containment
List<ToolUse> pool = new ArrayList<>(actual);
for (ToolUse e : expected) if (!pool.remove(e)) return false;
return true;
}
}
}
final class Normalize {
/** Whole numbers become Long, others Double, maps sorted, so 3 equals 3.0. */
static Object value(Object v) {
if (v instanceof Number n) {
double d = n.doubleValue();
return (d == Math.rint(d) && !Double.isInfinite(d)) ? (Object) (long) d : (Object) d;
}
if (v instanceof Map<?, ?> m) return map(m);
if (v instanceof List<?> l) return l.stream().map(Normalize::value).toList();
return v;
}
static Map<String, Object> map(Map<?, ?> m) {
Map<String, Object> r = new TreeMap<>();
m.forEach((k, x) -> r.put(String.valueOf(k), value(x)));
return r;
}
}
Scoring the final response
The Python ADK's response_match_score is documented as ROUGE-1 similarity, meaning unigram overlap between the response and the reference, with a default threshold of 0.8. The Java version below computes unigram F1 with clipped counts. Tokenisation and stemming differ between implementations, so do not assume it equals the Python score, and set your own threshold.
static double unigramF1(String reference, String candidate) {
Map<String, Integer> ref = counts(reference), cand = counts(candidate);
int overlap = 0;
for (var e : cand.entrySet()) overlap += Math.min(e.getValue(), ref.getOrDefault(e.getKey(), 0));
if (overlap == 0) return 0.0;
double precision = (double) overlap / total(cand), recall = (double) overlap / total(ref);
return 2 * precision * recall / (precision + recall);
}
static Map<String, Integer> counts(String s) {
Map<String, Integer> m = new HashMap<>();
for (String tok : s.toLowerCase(Locale.ROOT).split("[^a-z0-9]+")) {
if (!tok.isEmpty()) m.merge(tok, 1, Integer::sum);
}
return m;
}
static int total(Map<String, Integer> m) {
return m.values().stream().mapToInt(Integer::intValue).sum();
}
Worked example: one case, three scores
Take the case above with one more turn: the user says "Turn that one off." and the expected call is set_device_info with device_id device_1 and status OFF. The agent under test first calls get_user_prefs with no arguments, then list_devices with status ON, and replies "Device device_1 in the Living Room is currently on." On turn two it calls set_device_info with the same two arguments in the opposite key order.
| Check | Turn 1 | Turn 2 | Case score |
|---|---|---|---|
| Trajectory, EXACT | 0 (extra call) | 1 (key order is irrelevant) | 0.5, fails at 1.0 |
| Trajectory, IN_ORDER | 1 | 1 | 1.0, passes |
| Response F1 | 0.889 | scored separately | passes at 0.8 |
By hand: the reference tokenises to eight words: the, living, room, device, device, 1, is, on. The response gives ten, and all eight reference words are among them, so precision is 8/10 = 0.8, recall is 8/8 = 1.0 and F1 is 2 x 0.8 x 1.0 / 1.8 = 0.889.
Now the warning. A response of "The Living Room device (device_1) is off." shares seven of its eight words with the reference, so precision and recall are both 0.875 and it also passes at 0.8, while being factually wrong. Lexical overlap measures phrasing, not truth. Treat it as a tripwire, rely on the trajectory for correct lookups, and add a rubric judge for meaning; this site's LLM-as-judge scorer for ADK Java builds one on LlmAgent with a structured verdict.
Determinism: scripted models and repeats
Test the harness and scorers deterministically: subclass BaseLlm, whose abstract methods are generateContent(LlmRequest, boolean) returning Flowable<LlmResponse> and connect(LlmRequest), and feed it canned responses. The custom LLM guide covers that class in detail.
final class ScriptedLlm extends BaseLlm {
private final Deque<LlmResponse> script;
ScriptedLlm(List<LlmResponse> responses) {
super("scripted");
this.script = new ArrayDeque<>(responses);
}
@Override public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
LlmResponse next = script.poll();
return next == null ? Flowable.error(new IllegalStateException("script exhausted"))
: Flowable.just(next);
}
@Override public BaseLlmConnection connect(LlmRequest request) {
throw new UnsupportedOperationException("live mode not scripted");
}
}With a real model, even at temperature zero, one run is a sample, not a measurement. Run each case three to five times and decide in advance whether every repeat must pass or a majority suffices. Keep every repeat's trajectory: a case passing 4 of 5 times usually has one alternative path worth studying.
Wiring it into JUnit and CI
A JUnit 5 dynamic test per case gives per-case reporting for free. Keep a scripted suite for every commit and a live suite, run on merges or nightly, that holds the release gate.
class HomeAgentEvalTest {
static final int REPEATS = 3;
static final double F1_MIN = 0.6; // set from your own baseline, not copied from Python
@TestFactory
Stream<DynamicTest> evalSet() throws IOException {
return EvalSetLoader.load(Path.of("src/test/resources/evals/home.evalset.json")).stream()
.map(c -> DynamicTest.dynamicTest(c.id(), () -> {
for (int r = 0; r < REPEATS; r++) {
List<Invocation> got = Harness.replay(HomeAgent.build(), c);
for (int i = 0; i < got.size(); i++) {
Invocation inv = got.get(i);
Turn want = c.turns().get(i);
assertNull(inv.error(), () -> c.id() + " errored: " + inv.error());
assertTrue(Scorers.trajectoryMatches(want.expectedTools(), inv.tools(), MatchType.IN_ORDER),
() -> c.id() + " turn " + want.id() + " tools " + inv.tools());
assertTrue(Scorers.unigramF1(want.expectedText(), inv.finalText()) >= F1_MIN,
() -> c.id() + " turn " + want.id() + " said: " + inv.finalText());
}
}
}));
}
}Also write a JSON report with each case and repeat's tool calls, text, scores, latency, model and agent version; it is what you diff when the gate fails. Treat latency and tool-call counts as budgets: five calls where there used to be two is a regression even when the answer is right. The production counterpart of those numbers is covered in tool observability and metrics, and guardrail logic you may want to exercise from eval cases lives in ADK Java callbacks.
Failure modes
- State leaking between cases. Results depend on case order. Use one runner per case.
- Number types in arguments. Integer against Double breaks EXACT matching on every numeric argument. Normalise both sides.
- Errors counted as failures, or as passes. A quota error is neither. Report errors separately and fail the gate on the error rate.
- Over-specified trajectories. EXACT on an agent that looks things up fails every release. Prefer IN_ORDER.
- Trusting lexical scores. F1 passes wrong answers that reuse the reference's words, as the worked example shows. Pair it with trajectory checks and a judge.
- Expected outputs that rot. Changing tool data fails cases for the wrong reason. Use fixed tool fixtures.
- One run per case. A single pass hides a flaky path.
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| Python eval-set format | Shared cases, likely forward compatibility | Nested JSON, fields you do not use yet |
| Scripted model | Fast, free, deterministic | Tests the harness, not the agent's judgment |
| Live model in the gate | Measures real behaviour | Cost, latency, flakiness, provider drift |
| Strict match types | Catch procedure changes | False failures on harmless variation |
What to do next
- Write ten eval cases from real conversations in the Python eval-set format.
- Implement the loader, harness and scorers, and unit-test them with the scripted model, including numeric arguments.
- Choose a match type per case, defaulting to IN_ORDER.
- Run the live suite five times per case to measure your baseline, then set the F1 threshold and the repeat rule from that data.
- Add a rubric judge for the criteria that matter most, and report errors separately from failures.
- Run the live suite in CI on merges, publish the JSON report, and check new ADK Java releases for native evaluation.