An agent that calls a language model fails in more ways than a REST client. The network drops, the quota runs out, the request is too large, the API key is wrong, the model refuses, the answer stops halfway through a JSON object, or the model emits a function call that does not parse. Each of these needs a different response: retry, back off, fail fast and page someone, show the user a safe message, or ask the model again with a smaller budget. Treat them all as one generic exception and you either retry things that can never succeed or swallow outages that should wake someone up.
This article builds an error taxonomy for the Agent Development Kit for Java. It starts from where errors actually surface in ADK's model call path, defines a small set of classes with one handling rule each, gives a classifier and callback wiring you can drop into an LlmAgent, and covers streaming, testing and the dashboards that make the taxonomy useful. Retry mechanics themselves are covered in the companion article on retry layers; here the focus is naming the failure correctly so the right policy runs.
Where model errors surface in ADK Java
In ADK Java a model is a BaseLlm, and the agent flow calls generateContent(LlmRequest request, boolean stream), which returns a Flowable<LlmResponse>. A failure can arrive through two separate channels.
Thrown errors. The Flowable terminates with an exception. With the built-in Gemini model the Google GenAI Java SDK raises com.google.genai.errors.ApiException for non-2xx HTTP responses, with the subclasses ClientException for 4xx and ServerException for 5xx. The exception exposes code(), status() and message(). Network failures surface as I/O exceptions and timeouts, usually wrapped by the reactive operators, so the useful type is often a few levels down the cause chain. ADK itself throws LlmCallsLimitExceededException when an invocation exceeds RunConfig.maxLlmCalls, which defaults to 500.
In the flow, BaseLlmFlow attaches onErrorResumeNext to the model call. It runs plugin error callbacks and then the agent's onModelErrorCallback list. If one returns a response, that response replaces the error; if all return empty, the original exception is re-raised and reaches whoever subscribed to Runner.runAsync.
In-band failures. The HTTP call succeeds and a perfectly ordinary LlmResponse is emitted, but it carries bad news. LlmResponse has errorCode() and finishReason(), both Optional<FinishReason>, plus errorMessage(). For a Gemini response with a candidate, ADK always copies the candidate's finish reason; if the candidate has no parts and did not stop normally, it also copies that reason into errorCode and the finish message into errorMessage. With no candidate at all, errorCode comes from the prompt feedback's block reason, or is the sentinel Unknown error.. None of these are exceptions, so they never reach onModelErrorCallback; only afterModelCallback sees them.
The taxonomy: classes and handling rules
A taxonomy is only useful if each class maps to exactly one handling rule. The set below is small enough to put on a dashboard and specific enough that nobody argues about what to do. Status strings in the second column are the canonical Google API statuses that accompany those HTTP codes.
| Class | Signal | Retry? | Handling rule |
|---|---|---|---|
| TRANSPORT | I/O exception in the cause chain | Yes, bounded | Retry with backoff in one layer; degrade after the budget |
| TIMEOUT | TimeoutException, socket timeout | Once, if no output streamed | Retry with the remaining deadline; never after partial output |
| RATE_LIMITED | 429, RESOURCE_EXHAUSTED | Yes, slower | Back off with jitter; shed load; alert on sustained rate |
| SERVER | 500, 503, 504 | Yes, bounded | Retry; fail over to another model or region if configured |
| BAD_REQUEST | 400, INVALID_ARGUMENT | No | Fix the request: context too long, bad schema, bad parameter |
| AUTH | 401, 403 | No | Page: credentials or permissions are broken for everyone |
| NOT_FOUND | 404 | No | Page: model name or endpoint is wrong, often after a deprecation |
| PROMPT_BLOCKED | errorCode from prompt feedback | No | Safe user message; log for policy review |
| OUTPUT_BLOCKED | finish SAFETY, RECITATION, BLOCKLIST, PROHIBITED_CONTENT, SPII | No | Safe user message; do not resend the same prompt |
| TRUNCATED | finish MAX_TOKENS | Changed request | Continue or regenerate with a larger budget or smaller task |
| TOOL_PROTOCOL | finish MALFORMED_FUNCTION_CALL, UNEXPECTED_TOOL_CALL, TOO_MANY_TOOL_CALLS | Once, changed | Repair prompt or tool schema; cap the loop |
| CALL_BUDGET | LlmCallsLimitExceededException | No | Stop the invocation; investigate the agent loop |
| UNKNOWN | anything else, including the sentinel | No | Log raw details; review weekly and promote to a class |
Unknown errors are deliberately not retried, since retrying an unclassified failure hides it. The classes also separate faults that are yours (BAD_REQUEST, AUTH, NOT_FOUND, CALL_BUDGET) from faults that are the provider's (SERVER, RATE_LIMITED) and outcomes that are policy (the BLOCKED classes), because those go to different people.
A classifier for both channels
The classifier has two entry points, one per channel. The thrown path walks the cause chain because reactive operators wrap exceptions. The in-band path uses FinishReason.knownEnum(), which maps the string to the SDK's FinishReason.Known enum and returns FINISH_REASON_UNSPECIFIED for anything it does not recognise, which is why the raw errorMessage() must be kept alongside the class.
import com.google.adk.models.LlmCallsLimitExceededException;
import com.google.adk.models.LlmResponse;
import com.google.genai.errors.ApiException;
import com.google.genai.types.FinishReason;
import java.io.IOException;
import java.net.SocketTimeoutException;
import java.util.Optional;
import java.util.concurrent.TimeoutException;
public final class LlmErrors {
public enum LlmErrorClass {
TRANSPORT, TIMEOUT, RATE_LIMITED, SERVER, BAD_REQUEST, AUTH, NOT_FOUND,
PROMPT_BLOCKED, OUTPUT_BLOCKED, TRUNCATED, TOOL_PROTOCOL, CALL_BUDGET, UNKNOWN
}
/** Thrown channel: walk the cause chain, most specific type first. */
public static LlmErrorClass classify(Throwable t) {
for (Throwable c = t; c != null; c = c.getCause()) {
if (c instanceof LlmCallsLimitExceededException) return LlmErrorClass.CALL_BUDGET;
if (c instanceof ApiException api) return fromStatus(api.code());
if (c instanceof TimeoutException || c instanceof SocketTimeoutException) return LlmErrorClass.TIMEOUT;
if (c instanceof IOException) return LlmErrorClass.TRANSPORT;
}
return LlmErrorClass.UNKNOWN;
}
static LlmErrorClass fromStatus(int code) {
return switch (code) {
case 400 -> LlmErrorClass.BAD_REQUEST;
case 401, 403 -> LlmErrorClass.AUTH;
case 404 -> LlmErrorClass.NOT_FOUND;
case 429 -> LlmErrorClass.RATE_LIMITED;
default -> code >= 500 && code < 600 ? LlmErrorClass.SERVER : LlmErrorClass.UNKNOWN;
};
}
/** In-band channel: empty means the response is healthy. */
public static Optional<LlmErrorClass> classify(LlmResponse r) {
if (r.finishReason().isPresent()) { // a candidate came back: its reason wins
return fromFinish(r.finishReason().get().knownEnum());
}
return r.errorCode().map(code -> // no candidate: prompt block or sentinel
code.knownEnum() == FinishReason.Known.FINISH_REASON_UNSPECIFIED
? LlmErrorClass.UNKNOWN : LlmErrorClass.PROMPT_BLOCKED);
}
static Optional<LlmErrorClass> fromFinish(FinishReason.Known k) {
return switch (k) {
case STOP, FINISH_REASON_UNSPECIFIED -> Optional.empty();
case MAX_TOKENS -> Optional.of(LlmErrorClass.TRUNCATED);
case SAFETY, RECITATION, BLOCKLIST, PROHIBITED_CONTENT, SPII -> Optional.of(LlmErrorClass.OUTPUT_BLOCKED);
case MALFORMED_FUNCTION_CALL, UNEXPECTED_TOOL_CALL, TOO_MANY_TOOL_CALLS ->
Optional.of(LlmErrorClass.TOOL_PROTOCOL);
default -> Optional.of(LlmErrorClass.UNKNOWN);
};
}
}The order matters. A SAFETY stop with no text has both fields set to SAFETY, and checking errorCode first would misfile it as a blocked prompt. Only when no candidate came back does errorCode mean the prompt itself was blocked. That reflects how ADK's own response builder fills the fields for Gemini. A custom BaseLlm for another provider decides for itself what to put there, so if you wrap a different model, extend the classifier with that adapter's conventions instead of assuming Gemini's. Never classify by searching message text: messages change without notice and are localised.
Wiring it into callbacks
Wire the classifier into both callbacks. The callbacks record the class in session state, emit a metric, and decide whether to substitute a response. Returning Maybe.empty() from either callback means "leave it alone": the original response passes through, or the original exception is re-raised.
import com.google.adk.agents.CallbackContext;
import com.google.adk.agents.LlmAgent;
import com.google.adk.models.LlmRequest;
import com.google.adk.models.LlmResponse;
import com.google.genai.types.Content;
import com.google.genai.types.Part;
import io.reactivex.rxjava3.core.Maybe;
import java.util.List;
final class ErrorTaxonomyCallbacks {
static final String DEGRADED = "The assistant is busy right now. Please try again in a minute.";
static final String REFUSED = "I can't help with that request. A person can follow up if needed.";
static Maybe<LlmResponse> afterModel(CallbackContext ctx, LlmResponse resp) {
return LlmErrors.classify(resp).map(cls -> {
record(ctx, cls, resp.errorMessage().orElse(""));
return switch (cls) {
case PROMPT_BLOCKED, OUTPUT_BLOCKED -> Maybe.just(text(REFUSED));
default -> Maybe.<LlmResponse>empty(); // TRUNCATED etc: let the caller decide
};
}).orElse(Maybe.empty());
}
static Maybe<LlmResponse> onModelError(CallbackContext ctx, LlmRequest req, Exception e) {
LlmErrors.LlmErrorClass cls = LlmErrors.classify(e);
record(ctx, cls, String.valueOf(e.getMessage()));
return switch (cls) {
case RATE_LIMITED, SERVER, TRANSPORT, TIMEOUT -> Maybe.just(text(DEGRADED)); // retries already spent
default -> Maybe.empty(); // AUTH, NOT_FOUND, BAD_REQUEST: re-raise and page
};
}
static void record(CallbackContext ctx, LlmErrors.LlmErrorClass cls, String raw) {
ctx.state().put("llm_error_class", cls.name());
Metrics.increment("llm_errors_total", "class", cls.name()); // your metrics facade
Log.warn("llm error class={} raw={}", cls, raw);
}
static LlmResponse text(String s) {
return LlmResponse.builder()
.content(Content.builder().role("model").parts(List.of(Part.fromText(s))).build())
.turnComplete(true)
.build();
}
}
LlmAgent agent = LlmAgent.builder()
.name("support_agent")
.model(config.modelName()) // from config, never hard-coded
.instruction("Answer billing questions using the tools provided.")
.afterModelCallback(ErrorTaxonomyCallbacks::afterModel)
.onModelErrorCallback(ErrorTaxonomyCallbacks::onModelError)
.build();Metrics and Log stand in for whatever facades your service already uses. Note the ordering: a retry decorator wrapped around the model sits inside generateContent, so by the time onModelErrorCallback runs, transient errors have already used their retry budget. That is why the callback degrades rather than retrying again, and why configuration errors are deliberately re-raised: a friendly message for a revoked API key would turn a page-worthy outage into a silent one.
Worked example: one bad morning
Consider a billing support agent on a bad Monday morning:
- A quota spike. A marketing email drives traffic up fivefold. The API returns 429 with status RESOURCE_EXHAUSTED. The decorator backs off and retries twice, then gives up;
onModelErrorclassifies RATE_LIMITED and returns the degraded message. The dashboard shows RATE_LIMITED climbing from near zero to 4 percent of calls, which pages the on-call engineer to raise quota or shed load. Users see a polite delay message, not a stack trace. - A truncated tool argument. A long refund explanation hits the output token limit. The response arrives with finish reason MAX_TOKENS and half a sentence.
afterModelrecords TRUNCATED and leaves the response alone; the application layer, which knows the task, asks for a shorter summary. If the agent used an output schema, TRUNCATED is the class that explains the JSON parse failure that follows. - A safety stop. A user pastes an abusive message. The candidate finishes with SAFETY. The callback substitutes the refusal text and records OUTPUT_BLOCKED, which goes to the trust and safety review queue instead of the reliability dashboard.
- A deprecated model. Someone changes a config to a model name that is no longer served. Every call returns 404. The callback returns empty, the exception reaches the runner, the error rate alert fires within minutes, and the NOT_FOUND label tells the responder exactly where to look.
Streaming and partial output
Streaming changes two things. First, the finish reason arrives on the final chunk, after partial text has been emitted with partial() set. A TRUNCATED or OUTPUT_BLOCKED classification therefore happens after the user has seen some output; your UI should be able to mark a message as cut off or withdrawn rather than pretending it completed. Second, never retry transparently after an exception mid-stream: the user would see the first half twice and an emitted tool call could run twice. Track whether any chunk was forwarded and treat later failures as non-retryable.
Failure modes
Common ways a taxonomy goes wrong:
- Checking only exceptions. Half the failures are in-band. If your only handler is
onModelErrorCallback, SAFETY stops and truncations look like successes. - Treating empty content as success. A response with no parts and no finish reason is a failure. Classify it as UNKNOWN and keep the raw message.
- Swallowing configuration errors. A callback that returns a friendly message for every class hides AUTH and NOT_FOUND from alerting.
- Retrying in the callback. Retry belongs in one layer, normally a model decorator. A second retry in the callback multiplies attempts during exactly the overload you are trying to ride out.
Testing the taxonomy
The classifier is pure, so test it with a table. Build each input with the real types rather than mocks, so a change in the SDK breaks the test instead of production.
@ParameterizedTest
@MethodSource("cases")
void classifies(Object input, LlmErrors.LlmErrorClass expected) {
LlmErrors.LlmErrorClass actual = input instanceof Throwable t
? LlmErrors.classify(t)
: LlmErrors.classify((LlmResponse) input).orElse(null);
assertEquals(expected, actual);
}
static Stream<Arguments> cases() {
return Stream.of(
arguments(new ApiException(429, "RESOURCE_EXHAUSTED", "quota"), RATE_LIMITED),
arguments(new RuntimeException(new ApiException(503, "UNAVAILABLE", "x")), SERVER),
arguments(new RuntimeException(new SocketTimeoutException()), TIMEOUT),
arguments(LlmResponse.builder().finishReason(new FinishReason("MAX_TOKENS")).build(), TRUNCATED),
arguments(LlmResponse.builder().finishReason(new FinishReason("SAFETY"))
.errorCode(new FinishReason("SAFETY")).build(), OUTPUT_BLOCKED),
arguments(LlmResponse.builder().errorCode(new FinishReason("PROHIBITED_CONTENT")).build(), PROMPT_BLOCKED),
arguments(LlmResponse.builder().errorCode(new FinishReason("Unknown error.")).build(), UNKNOWN));
}Then add one integration test per channel: a scripted fake BaseLlm run through a real LlmAgent, asserting on the runner's event and the state key the callback wrote.
Operating it: metrics and alerts
Make the taxonomy visible. Count every model call and every classified failure by class and model name, and build three views. Reliability shows TRANSPORT, TIMEOUT, RATE_LIMITED and SERVER as a rate against total calls; this is the view tied to your availability objective. Correctness shows BAD_REQUEST, AUTH, NOT_FOUND and CALL_BUDGET, any non-zero value of which is a bug or misconfiguration. Policy shows the two BLOCKED classes and TRUNCATED; their trends tell product and safety owners how prompts, users and limits are interacting.
Exclude user-caused classes such as OUTPUT_BLOCKED from the availability objective, and review the UNKNOWN bucket weekly, promoting recurring patterns to named classes.
Related reading
For the retry side of this design read retry layers for ADK Java model calls, and for recovery patterns at the agent level see error recovery in ADK Java. The callback mechanics are covered in ADK Java callbacks, the model contract in the BaseLlm interface, tool-side failures in tool error wrapping, and fakes for tests in testing custom LLMs.
What to do next
- List every place your agents handle model failures today, and check whether any of them reads
finishReason()orerrorCode(). - Add the classifier and the two callbacks, starting with record-only behaviour and no substitutions.
- Ship the per-class metric and build the reliability, correctness and policy views.
- Turn on substitutions for the BLOCKED classes and the degraded message for transient classes once the numbers look sane.
- Make sure AUTH and NOT_FOUND still reach the runner and page someone; test it by pointing a staging agent at a bad model name.
- Schedule a weekly review of UNKNOWN and keep the table-driven test in step with the classes.