Most traffic to an assistant is easy: order status, opening hours, a definition, a yes or no. A smaller share needs real reasoning: comparing two plans, debugging a stack trace, drafting a policy exception. Sending everything to your strongest model pays premium cost and latency on the easy majority; sending everything to the fastest model fails the hard minority. A query router decides, per turn, which model tier handles the request.
The agent router article covers the other routing axis, which specialist agent owns a turn, and deliberately sets model tiers aside. This article builds the tier router for ADK Java as a BaseLlm decorator: where it plugs in, the one request field it must rewrite, how to keep a turn on one model while tools run, how to classify, how to escalate safely, and how to prove the router is not quietly degrading answers.
When a tier router pays off
A tier router is worth building when three things hold. Your traffic is skewed toward simple requests. The price or latency gap between tiers is large; across current model families it is commonly several-fold, but measure your own. And you can tell easy from hard cheaply, before the call, from the request alone.
The router is not free. Every misroute downward produces a worse answer that a user sees; every misroute upward wastes money. The two errors are not symmetric: a downward misroute on a refund question costs trust, while an upward misroute on a greeting costs a fraction of a cent. Design the classifier to be conservative, sending anything uncertain to the strong tier, and measure both error rates rather than only the cost saving.
Where the router plugs in
ADK Java gives you three places to choose a model, and they are not equivalent.
| Option | Granularity | Verdict |
|---|---|---|
| Fixed model per LlmAgent | Per agent | Right when one agent always needs one tier; no per-turn routing. |
| Before-model callback | Per model call | Receives an LlmRequest.Builder, so it can set the model name on one Gemini client; it cannot switch providers, escalate on error or route live sessions. |
| BaseLlm decorator | Per model call | Owns the call, picks a delegate, rewrites the request; testable in isolation. Use this. |
The decorator pattern is covered in general in the BaseLlm interface article. BaseLlm is an abstract class with a model name, generateContent(LlmRequest, boolean stream) returning Flowable<LlmResponse>, and connect(LlmRequest) for live sessions. LlmAgent.builder().model(...) accepts a BaseLlm instance directly, so the agent never knows a router is there.
The architecture
The LlmAgent is unchanged: same instruction, same tools, same callbacks. RoutingLlm receives each assembled LlmRequest. The QueryClassifier sees only the latest user text and returns FAST or STRONG. The two delegates are ordinary BaseLlm instances, for example two Gemini clients built with different model names. The route log records every decision so misroutes can be found later.
One detail is easy to miss. The Gemini implementation in ADK Java resolves the model it calls with llmRequest.model().orElse(model()): a model name on the request wins over the client's own. If the request arrives carrying a model name, delegating it unchanged can send the call to the wrong tier however carefully you chose the delegate. The router therefore rebuilds the request with request.toBuilder().model(target.model()).build() before delegating.
Keeping a turn on one tier
One user turn can produce several model calls. The model emits a function call, the runtime runs the tool, appends the function response to the contents and calls the model again. If the router re-decided on each call using the whole growing request, a turn could start on the strong model, plan a tool call, and then have the tool result interpreted by the fast model, which never saw its own reasoning. Mixed-tier turns are hard to debug and can break model-specific conventions in multi-step tool use.
The fix is to make the decision a pure function of the latest user text. Function responses are added to the contents as parts without text, so scanning backwards for the most recent user content with a text part finds the same message on every call within the turn, and the classifier returns the same tier. No session state, no cache, nothing to expire. The rule this imposes: the classifier must not read conversation length, tool results or anything else that changes inside a turn.
The RoutingLlm decorator
The router in full. It assumes RxJava 3, as ADK Java uses, and the google-genai types for Content and Part, whose accessors return Optional.
public final class RoutingLlm extends BaseLlm {
public enum Tier { FAST, STRONG }
public interface QueryClassifier {
Tier classify(String latestUserText); // must depend on the text only
}
public interface RouteLog {
void decided(Tier tier, String reason, int textLength);
void escalated(Throwable cause);
}
private final BaseLlm fast;
private final BaseLlm strong;
private final QueryClassifier classifier;
private final RouteLog log;
public RoutingLlm(BaseLlm fast, BaseLlm strong, QueryClassifier classifier, RouteLog log) {
super(strong.model()); // the name the runtime sees; tiers live in the log
this.fast = fast;
this.strong = strong;
this.classifier = classifier;
this.log = log;
}
@Override
public Flowable<LlmResponse> generateContent(LlmRequest request, boolean stream) {
return Flowable.defer(() -> {
String text = latestUserText(request.contents());
Tier tier;
try {
tier = text.isEmpty() ? Tier.STRONG : classifier.classify(text);
} catch (RuntimeException e) {
tier = Tier.STRONG; // a broken classifier fails upward
}
log.decided(tier, text.isEmpty() ? "no-text" : "classifier", text.length());
if (tier == Tier.STRONG) {
return call(strong, request, stream);
}
AtomicBoolean emitted = new AtomicBoolean(false);
return call(fast, request, stream)
.doOnNext(r -> emitted.set(true))
.onErrorResumeNext(err -> {
if (emitted.get()) {
return Flowable.error(err); // never splice two models into one answer
}
log.escalated(err);
return call(strong, request, stream);
});
});
}
@Override
public BaseLlmConnection connect(LlmRequest request) {
return strong.connect(request.toBuilder().model(strong.model()).build());
}
private static Flowable<LlmResponse> call(BaseLlm target, LlmRequest request, boolean stream) {
return target.generateContent(request.toBuilder().model(target.model()).build(), stream);
}
static String latestUserText(List<Content> contents) {
for (int i = contents.size() - 1; i >= 0; i--) {
Content content = contents.get(i);
if (!"user".equals(content.role().orElse(""))) {
continue;
}
String text = content.parts().orElse(List.of()).stream()
.map(part -> part.text().orElse(""))
.collect(Collectors.joining(" "))
.strip();
if (!text.isEmpty()) {
return text; // function responses carry no text: skipped
}
}
return "";
}
}Wiring it in takes one line on the agent. Model names come from configuration so a tier can be swapped without a code change.
BaseLlm fast = Gemini.builder().modelName(cfg.fastModel())
.apiKey(System.getenv("GOOGLE_API_KEY")).build();
BaseLlm strong = Gemini.builder().modelName(cfg.strongModel())
.apiKey(System.getenv("GOOGLE_API_KEY")).build();
LlmAgent agent = LlmAgent.builder()
.name("support")
.model(new RoutingLlm(fast, strong, new RuleClassifier(400), routeLog))
.instruction("Answer customer questions. Use lookup_order for order status.")
.tools(lookupOrder)
.build();Live sessions go straight to the strong tier: a bidirectional audio session cannot be re-routed mid-stream, so pick one model and keep it. Escalation happens only when the fast call fails before emitting anything; once a partial response has reached the caller, switching models would produce an answer stitched from two sources, so the error propagates instead and your normal retry policy takes over.
Classifying queries
Start with rules. They are free, deterministic, explainable and easy to test. A useful first classifier sends a request to the fast tier only if it is short, contains no code, and matches none of a list of hard-intent words drawn from your own traffic.
public final class RuleClassifier implements RoutingLlm.QueryClassifier {
private static final Pattern HARD = Pattern.compile(
"\\b(why|compare|explain|debug|design|trade-?offs?|refund|cancel|legal|contract|complain\\w*)\\b",
Pattern.CASE_INSENSITIVE);
private final int maxFastChars;
public RuleClassifier(int maxFastChars) { this.maxFastChars = maxFastChars; }
@Override
public RoutingLlm.Tier classify(String text) {
if (text.length() > maxFastChars) return RoutingLlm.Tier.STRONG;
if (text.contains("```") || text.contains("Exception")) return RoutingLlm.Tier.STRONG;
if (HARD.matcher(text).find()) return RoutingLlm.Tier.STRONG;
return RoutingLlm.Tier.FAST;
}
}When rules plateau, move to a learned classifier: label a few thousand real turns as fast-ok or needs-strong, using the strong model's answer as the reference and a judge or a human to grade the fast model's answer, then train a small model on embeddings of the user text. It runs in milliseconds in-process. Calling an LLM to classify is the third option and usually the wrong one: it adds a full model round trip to every turn and must itself be cheap enough not to erase the saving.
Whatever you use, put a threshold on its confidence and send low-confidence cases to the strong tier. The threshold, not the model, is the main dial between cost and quality.
Worked example: a support assistant
Take a support assistant handling 20,000 turns a day. Use illustrative relative costs: a fast call costs 1 unit and a strong call 12 units per turn. Everything on the strong tier costs 240,000 units a day.
Suppose a hypothetical offline evaluation on 500 labelled turns finds that the fast tier's answer is acceptable for 355 of them (71 percent). The rule classifier sends 310 turns (62 percent) to the fast tier. Of those, 291 were acceptable and 19 were not: a downward misroute rate of 6.1 percent of fast-routed traffic, 3.8 percent of all turns. Meanwhile 64 turns that the fast tier could have handled went to the strong tier, the upward misroutes.
Projected cost: 0.62 × 1 + 0.38 × 12 = 5.18 units per turn, plus escalations. At a 1 percent fast-tier error rate, escalations add 0.0062 × 12 ≈ 0.07 units, for about 5.25 units per turn or 105,000 units a day, a 56 percent saving. Whether that is a good trade depends entirely on the 19 bad answers: read them. If they cluster on one intent, such as billing disputes, add that intent to the hard list and re-run; the saving drops a little and the misroute rate drops a lot. Ship only when the downward misroute rate on high-stakes intents is near zero.
Observability and shadow mode
Log, for every model call: the tier, the reason, the text length, latency, token usage from the response's usage metadata, and whether an escalation happened. Then watch four numbers: the fast-tier share, the escalation rate, user feedback or resolution rate split by tier, and cost per turn. A sudden rise in fast-tier share after a release usually means the classifier input changed, for example a prompt template now prepends text to the user message. Counting tokens across models explains why per-tier usage must be normalised before you compare costs.
Before switching on routing, run it in shadow: classify every turn, log the decision, but keep serving from the strong tier. A week of shadow logs gives you the real traffic mix and a sample of would-be fast turns to grade, at no risk to users.
Failure modes
- Request model not rewritten. The delegate calls whatever name the request carries; the router logs FAST while the bill says otherwise. Assert the model on a captured request in a unit test.
- Mid-turn tier flips. A classifier that reads contents size or tool output changes its answer between calls within one turn. Keep it a function of the latest user text.
- Tool schema mismatch. Both tiers receive the same function declarations; a fast model that handles them poorly produces malformed calls. Evaluate tool use, not just text answers, on the fast tier.
- Escalation after partial output. Splicing a strong answer onto a partial fast answer duplicates or contradicts text. Escalate only before the first emission.
- Silent quality drift. A model upgrade on either tier changes the acceptable-on-fast set. Re-run the labelled evaluation whenever a tier's model changes.
- Classifier gaming. Users learn that a long or code-heavy message gets the better model. That is usually acceptable; just make sure the router never controls permissions or tools.
Trade-offs
Rules are transparent and free but plateau quickly. A learned classifier captures more of the saving but needs labelled data and retraining. A cascade that always tries the fast model first and escalates on a quality check catches more misroutes, but it needs a reliable verifier and pays fast-model latency on every hard query. Routing also adds a dimension to everything you test: callbacks, guardrails and evaluations must pass on both tiers. If your traffic is mostly hard, or the tier gap is small, a single model is simpler and the right answer.
Keep the two routing axes separate. Choose the specialist with the agent router, choose the tier inside each specialist with this decorator, and see model call orchestration for how both fit into one turn.
What to do next
- Export a week of real user turns and label a few hundred as fast-ok or needs-strong.
- Implement RoutingLlm and a unit test that captures the delegated LlmRequest and asserts its model name for each tier.
- Add a test turn with a tool call and assert that every model call in the turn hit the same tier.
- Write a RuleClassifier from your own hard intents and measure both misroute rates offline.
- Run the router in shadow mode for a week, logging decisions while serving from the strong tier.
- Enable routing for a small traffic percentage and compare resolution rate and cost per turn by tier.
- Re-run the labelled evaluation whenever either tier's model changes.