A smoke test answers one question quickly: is this build, in this environment, broken in an obvious way? It is not an evaluation of answer quality and not a full regression suite. For an ordinary service a smoke test is a health endpoint and one request. For an agent that is not enough, because the failures that matter are quiet: a tool whose credentials expired still produces a polite answer, and a renamed tool still leaves the model talking fluently about what it would have found.

This article builds a layered smoke suite for an ADK Java agent: a wiring layer that runs without a model, a dependency layer that probes each backend directly, and one live turn whose assertions read ADK events rather than the answer's prose. It covers where each layer runs, how to keep smoke traffic out of your data, and how to handle flakiness. The HTTP client that drives a deployed revision is shown in Agent Deploy Pipeline; the suite here is what that client should be checking.

What a smoke test is for an agent

Three properties separate smoke tests from the rest of your testing. They are fast: the whole suite finishes in under a minute, so it can gate a deploy or run every few minutes. They run against real infrastructure: the point is to find the misconfigured secret, the missing IAM grant or the unreachable MCP server, which integration tests with fakes cannot see. And they are deterministic in what they assert, even though the model is not.

That last property is the design problem. A live model call varies in wording, so a smoke test that compares text is flaky, and one that only checks for a 200 response passes when the agent is broken. The way out is that ADK records structure for every turn: which tools were called with which arguments, what each returned, whether the final event is a final response, whether the model reported an error code, and how many tokens it used. Those are stable enough to assert on.

Layer 1: wiring without a model

Three smoke layers, cheapest firstLayer 1: wiringno model, no network, secondsLayer 2: dependenciesreal backends, read-only probesLayer 3: live turnreal model, asserted on eventsagent tree buildsunique agent and tool namesevery tool declares itselfcredentials acceptedfixture records presentMCP servers list toolsexpected tool was calledno error-shaped responsesfinal response, in budgetStop at the first failing layer: a 401 in layer 2 never costs a model callRuns at build (1), pre-traffic after deploy (1, 2, 3) and on a schedule in production (2, 3).
Each layer is more expensive and more realistic than the one before; the suite stops at the first failure.

The first layer builds the production agent tree with production configuration and checks it without calling a model. It catches the largest class of deploy breakages, which are wiring mistakes: a sub-agent that was renamed, two tools with the same name, a tool missing its description, a toolset that cannot be resolved.

@Test
void productionAgentTreeIsWired() {
  BaseAgent root = Agents.root(Config.load("prod-smoke"));   // same factory as production
  Set<String> agentNames = new HashSet<>();
  Deque<BaseAgent> todo = new ArrayDeque<>(List.of(root));
  while (!todo.isEmpty()) {
    BaseAgent a = todo.pop();
    assertTrue(agentNames.add(a.name()), "duplicate agent name: " + a.name());
    todo.addAll(a.subAgents());
    if (a instanceof LlmAgent llm) {
      List<BaseTool> tools = llm.canonicalTools().toList().blockingGet();
      Set<String> toolNames = new HashSet<>();
      for (BaseTool t : tools) {
        assertTrue(toolNames.add(t.name()), a.name() + " has two tools named " + t.name());
        assertFalse(t.description() == null || t.description().isBlank(),
            t.name() + " has no description");
        t.declaration().ifPresent(d ->
            assertEquals(Optional.of(t.name()), d.name(), "declaration name differs"));
      }
      assertTrue(Expected.TOOLS.get(a.name()).equals(toolNames),
          a.name() + " tools changed: " + toolNames);
    }
  }
}

The final assertion compares each agent's tool names with a checked-in expectation. That turns an accidental removal into a failing test, and an intentional change into a one-line diff in review. Built-in tools such as Google Search may have no function declaration, which is why the declaration check is conditional.

One caution: canonicalTools() resolves toolsets, and resolving an MCP toolset means talking to its server. In a build with no network, substitute a configuration that points MCP toolsets at a local stub, and leave the real servers to layer 2.

Layer 2: dependency probes

The second layer checks each external dependency the tools rely on, directly, with the credentials the deployed agent will use, and without a model in the loop. Each probe is read-only and touches a known fixture so that it proves access, not just reachability.

public interface SmokeProbe {
  String name();
  Duration budget();
  void run() throws Exception;      // throw on failure, with a message an operator can act on
}

List<SmokeProbe> probes = List.of(
    probe("orders-db", Duration.ofSeconds(3),
        () -> orders.findById("ORD-SMOKE-0001").orElseThrow(() -> new AssertionError("fixture missing"))),
    probe("crm-api", Duration.ofSeconds(5),
        () -> crm.getCustomer("CUST-SMOKE").requireStatus(200)),
    probe("docs-mcp", Duration.ofSeconds(10),
        () -> require(docsToolset.getTools(readonlyContext).toList().blockingGet().size() > 0)),
    probe("secrets", Duration.ofSeconds(1),
        () -> Secrets.requireAll("CRM_TOKEN", "ORDERS_DB_URL")));

Map<String, String> failures = SmokeRunner.runAll(probes);   // parallel, each with its budget
if (!failures.isEmpty()) throw new SmokeFailure("dependency layer", failures);

Calling the backend client directly rather than through the tool keeps the probe simple, but it can miss a bug in how the tool builds its request. If a tool's core is a plain method, as recommended in Agent Unit Testing, call that method with the fixture id instead. Never probe with a write: a smoke suite that creates refunds or tickets will eventually do it in production at three in the morning.

Layer 3: one live turn, asserted on events

The third layer runs one real turn through the production runner wiring, with the real model and session service, as a dedicated smoke user. The prompt is chosen to force exactly one known tool call against the fixture, and the assertions read the event stream.

Runner runner = App.buildRunner(Config.load("prod"));        // same wiring as serving
String user = "smoke-" + System.getenv("REVISION");
Session s = runner.sessionService().createSession(runner.appName(), user).blockingGet();
try {
  List<Event> events = runner
      .runAsync(user, s.id(), Content.fromParts(Part.fromText(
          "Look up order ORD-SMOKE-0001 and tell me its status.")))
      .timeout(60, TimeUnit.SECONDS)
      .toList()
      .blockingGet();
  TurnVerdict.of(events)
      .expectToolCalled("lookup_order", args -> "ORD-SMOKE-0001".equals(args.get("order_id")))
      .expectNoErrorResponses()
      .expectFinalResponse()
      .expectTokensBelow(12_000)
      .assertPassed();
} finally {
  runner.sessionService().deleteSession(runner.appName(), user, s.id()).blockingAwait();
}
TurnVerdict expectNoErrorResponses() {
  for (Event e : events) {
    e.errorCode().ifPresent(code -> fail("model error " + code + ": " + e.errorMessage().orElse("")));
    for (FunctionResponse r : e.functionResponses()) {
      Map<String, Object> body = r.response().orElse(Map.of());
      if (body.containsKey("error") || "error".equals(body.get("status"))) {
        fail("tool " + r.name().orElse("?") + " returned " + body);
      }
    }
  }
  return this;
}

The error check is the heart of the suite, and it matches how ADK reports problems. When a FunctionTool body throws, ADK logs the exception and sends the model {"status": "error", "message": "An internal error occurred."}. When a tool requires confirmation, the first call returns an error entry asking for approval. In both cases the model goes on to write a reasonable apology, and a text check passes. The function response does not lie.

Pick a smoke prompt that cannot be answered without the tool, and avoid confirmation-gated tools. Alternatively, test the gate on purpose: expect an adk_request_confirmation function call event, assert it appeared, and never approve it. Keep exactly one live turn; answer quality belongs in evaluation, as in evals in CI.

Where each layer runs

WhereLayersWhyWatch out for
Build / pull request1Free and fast; catches wiring changes before mergeNo production secrets in CI; stub MCP
Container start1 (and secret presence)Fail before the instance takes trafficDo not call the model in a readiness probe: cost and rate limits multiply by replica count
After deploy, before traffic1, 2, 3Proves this revision works in this environmentRun against the new revision only, not the load-balanced URL
Scheduled in production2, 3Detects expiry, quota and drift between deploysRate: every 5 to 15 minutes is plenty; alert on two consecutive failures

The post-deploy run is the one that gates promotion; a failing layer stops the rollout before the canary sees any user. The scheduled run catches what no deploy caused: a rotated key, an exhausted quota, an expired OAuth grant, a model version change on the provider side.

Smoke hygiene

  • Identify smoke traffic. Use a smoke- user prefix and set it in trace attributes, so dashboards, billing reports and analytics can exclude it.
  • Keep it out of memory. If your application calls addSessionToMemory when sessions end, skip smoke users; otherwise the long-term store fills with fixture conversations.
  • Clean up. Delete the session in a finally block and sweep any smoke sessions older than an hour in case a run was killed.
  • Budget it. One turn costs a few thousand tokens. Every five minutes, across three regions, that is real money each month; size the schedule deliberately.
  • Retry narrowly. Retry once on transport errors, HTTP 429 and 5xx from the model. Never retry an assertion failure: a tool error that disappears on retry is a flaky dependency, which is exactly what the suite exists to report.

Worked example: a rename and a rotated token

A team ships a release that renames the order tool from lookup_order to get_order and, in the same week, rotates the CRM token. Layer 1 fails in the pull request: the tool-name expectation for support_agent no longer matches. The developer updates the expectation and notices the instruction still says "use lookup_order", so they fix that too. Without layer 1 the model would have improvised, and the HTTP smoke check for a status word would probably have passed.

After deploy, layer 2 fails in 300 milliseconds: crm-api returns 401 because the new revision mounts the old secret version. No model call was made, and the error names the probe. After the secret is fixed, layer 3 passes: one call to get_order with the fixture id, no error-shaped responses, a final response, 6,800 tokens.

Two weeks later the scheduled run fails twice in a row. Layer 2 is green, but layer 3 shows get_order returning {"status": "error", "message": "An internal error occurred."}. The application log has the stack trace: a schema migration added a non-null column the fixture row lacked. The final answer for both runs read "I'm having trouble retrieving that order right now", which a text check would have accepted.

Failure modes

  • Green suite, broken agent. Assertions on HTTP status or a keyword only. Assert on function calls, function responses and final-response events.
  • Flaky live turn. The prompt allows an answer without the tool. Make the fixture id the only route to the answer, and assert the tool call before anything else.
  • Smoke that writes. A probe or prompt creates real records. Use read-only operations on fixtures and review the suite like production code.
  • Fixture drift. Someone deletes ORD-SMOKE-0001 during a cleanup. Protect fixtures by naming convention and make layer 2 report a missing fixture distinctly from an access failure.
  • Testing the wrong revision. The suite runs against the load-balanced URL and exercises the old version. Target the new revision explicitly.
  • Slow suite. Layers run in sequence with generous timeouts and the gate takes five minutes. Run layer 2 probes in parallel with per-probe budgets.

Trade-offs

ChoiceGainsCosts
Live model in smoke vs scripted modelCatches model access, quota and real tool selectionTokens, latency, some nondeterminism
Probe clients vs tool coresSimple, independent of agent codeCan miss bugs in the tool's request building
One turn vs a short scenarioFast, stableMisses multi-turn and state bugs; leave those to integration tests
Scheduled smoke vs real-traffic alertsSignal even with no usersCost; synthetic traffic to exclude everywhere

What to do next

  1. Write the layer 1 test against your production agent factory, with a checked-in map of expected tool names per agent.
  2. Create read-only fixtures in each backend and a probe per dependency with its own budget.
  3. Choose one smoke prompt that needs exactly one tool, and assert on events: tool called with the fixture argument, no error-shaped function responses, a final response, a token ceiling.
  4. Wire layers 1 to 3 into the deploy so a failure stops promotion before any traffic shift.
  5. Schedule layers 2 and 3 every 5 to 15 minutes, alerting on two consecutive failures.
  6. Exclude smoke users from memory ingestion, analytics and billing, and sweep leftover sessions.
Key takeaway: An agent fails quietly, so its smoke suite has to look below the answer. Check the wiring without a model, probe each dependency directly with real credentials, and run one live turn whose assertions read ADK events: the expected tool call, no error-shaped function responses, a final response within a token budget. Gate deploys on all three layers and run the last two on a schedule.