An agent that answers correctly but cannot be used by a blind customer, a voice-only caller or someone with a cognitive disability has failed those users completely. Chat interfaces make this easy to get wrong: models love markdown tables, streamed tokens can flood a screen reader, and instructions like "press the green button" exclude anyone who cannot see colour. Accessibility in an agent spans three layers: what the model writes, how responses are adjusted for a particular user, and how the client presents them.
ADK for Java has no accessibility feature to switch on, and this guide does not pretend otherwise. It shows how to build accessibility from documented ADK pieces: rules in the agent instruction, a user-scoped profile in session state, an afterModelCallback that checks and deterministically repairs output, and a client that announces responses properly. It maps each piece to WCAG 2.2 success criteria and finishes with a worked example, tests and failure modes. Check API details against the ADK release you use.
What accessibility means for an agent
The Web Content Accessibility Guidelines (WCAG) 2.2 are written for web content, but their success criteria translate directly into requirements for conversational agents. The useful mapping:
| WCAG 2.2 criterion | What it means for an agent |
|---|---|
| 1.1.1 Non-text Content | Images, charts and files the agent returns need a text alternative that serves the same purpose. |
| 1.3.1 Info and Relationships | Structure conveyed visually (tables, indentation, layout) must also be available in text or markup. |
| 1.4.1 Use of Color | Never rely on colour alone, including in the agent's own wording. |
| 2.2.1 Timing Adjustable | Session or confirmation timeouts must be adjustable or extendable; screen reader users are slower. |
| 2.4.4 Link Purpose (In Context) | Link text says where it goes. |
| 3.3.1 Error Identification | When a request fails, say what went wrong and how to fix it, in text. |
| 4.1.3 Status Messages | "Working", "answer ready" and errors must reach assistive technology without moving focus. |
Reading level is criterion 3.1.5 at level AAA, which most organisations do not mandate, but plain language helps everyone, including people reading in a second language and anyone using the agent under stress. Treat it as a per-user preference with a sensible default.
Architecture: defaults, profiles and repair
The design principle is accessible by default, adjusted per user. The instruction asks every response to be readable by everyone: answer first, short sentences, descriptive links, no colour-only references. That costs sighted users nothing. Preferences that do have a cost for others, such as removing all markdown for a voice channel, live in a per-user profile and are applied by deterministic code after the model, never by hoping the model remembers.
Why repair in a callback instead of asking the model to follow the profile? Because instructions are probabilistic. A model told "no tables for this user" will still emit one occasionally, especially after a tool returns tabular data. A callback runs on every final model response, sees the profile, and applies transformations that are testable with unit tests. The model's job shrinks to writing well; the code's job is guaranteeing the format.
The accessibility profile in session state
ADK session state supports scoped keys: a user: prefix ties the value to the user id and shares it across all of that user's sessions, app: shares across the application and temp: lasts one invocation. A profile belongs under user: so a person sets it once. Seed it when you create the session from your own user settings store, and let users change it in conversation through a small tool if you want that.
// Stored under a user: key so it follows the user across sessions.
public record A11yProfile(
boolean screenReader, // avoid tables and ASCII art; describe structure in words
boolean voiceOutput, // no markdown at all; short sentences; spell out symbols
boolean plainLanguage, // shorter sentences, define terms, one idea per sentence
boolean reducedMotion) { // client hint: no animated typing effect
public static final String STATE_KEY = "user:a11y_profile";
public static A11yProfile from(Object raw) {
if (!(raw instanceof Map<?, ?> m)) return new A11yProfile(false, false, false, false);
return new A11yProfile(
Boolean.TRUE.equals(m.get("screen_reader")),
Boolean.TRUE.equals(m.get("voice_output")),
Boolean.TRUE.equals(m.get("plain_language")),
Boolean.TRUE.equals(m.get("reduced_motion")));
}
}Keep the profile to settings that change output. Do not store disability diagnoses: you need to know that someone uses a screen reader, not why, and collecting health information creates privacy obligations you do not want. Let users choose; do not infer from behaviour, because guesses are often wrong and feel intrusive.
Accessible defaults in the instruction and tools
The instruction carries the defaults. Each rule maps to a criterion above, and each is phrased as a concrete behaviour the model can follow rather than "be accessible":
public static final LlmAgent ROOT_AGENT = LlmAgent.builder()
.name("support_agent")
.model("<a model id your project can use>")
.instruction("""
You help customers with orders and billing.
Write so that every reader, including screen reader and voice users, can follow:
- Put the answer first, then the detail.
- Use sentences under 25 words. Define any term a customer may not know.
- Never refer to anything only by colour, position or shape ("the red button").
- Link text must say where it goes, never "click here" or a bare URL.
- Prefer short lists to tables. If you compare items, also state the result in words.
- When you describe an image the user sent, describe what matters to their question.
""")
.tools(FunctionTool.create(OrderTools.class, "getOrderStatus"))
.afterModelCallback(A11yCallback::apply)
.build();Tool results matter as much as instructions. If getOrderStatus returns raw codes such as SHP_EXC_07, the model will echo them. Return human-readable fields from tools (a state plus a sentence explaining it) so the model has accessible material to work with. The same applies to error results: a structured error with a user-facing message satisfies 3.3.1 far more reliably than a stack trace the model has to interpret.
Checking and repairing output in afterModelCallback
The callback below follows the documented Java shape: it receives the CallbackContext and the LlmResponse, returns Maybe.empty() to keep the model's response, or Maybe.just(...) to replace it. It skips partial streaming chunks and responses that contain function calls, reads the profile from ctx.state(), records violations, and rewrites only what can be rewritten deterministically:
public final class A11yCallback {
private static final Pattern TABLE_ROW = Pattern.compile("(?m)^\\s*\\|.*\\|\\s*$");
private static final Pattern VAGUE_LINK =
Pattern.compile("(?i)\\[(click here|here|this|link|read more)\\]\\(");
private static final Pattern COLOUR_ONLY =
Pattern.compile("(?i)\\bthe (red|green|blue|yellow|orange) (button|link|icon|box)\\b");
public static Maybe<LlmResponse> apply(CallbackContext ctx, LlmResponse resp) {
if (resp.partial().orElse(false)) return Maybe.empty(); // judge whole turns
List<Part> parts = resp.content().flatMap(Content::parts).orElse(List.of());
if (parts.stream().anyMatch(p -> p.functionCall().isPresent())) return Maybe.empty();
String text = parts.stream().map(p -> p.text().orElse(""))
.collect(Collectors.joining("\n"));
A11yProfile profile = A11yProfile.from(ctx.state().get(A11yProfile.STATE_KEY));
List<String> violations = new ArrayList<>();
if (VAGUE_LINK.matcher(text).find()) violations.add("vague_link_text");
if (COLOUR_ONLY.matcher(text).find()) violations.add("colour_only_reference");
boolean hasTable = TABLE_ROW.matcher(text).find();
if (hasTable && (profile.screenReader() || profile.voiceOutput())) violations.add("table");
A11yMetrics.record(ctx.agentName(), violations); // your metrics sink
String repaired = text;
if (hasTable && (profile.screenReader() || profile.voiceOutput())) {
repaired = Tables.toSentences(repaired); // deterministic; skips fenced code blocks
}
if (profile.voiceOutput()) repaired = Markdown.toSpeech(repaired);
if (repaired.equals(text)) return Maybe.empty(); // keep the original
return Maybe.just(LlmResponse.builder()
.content(Content.fromParts(Part.fromText(repaired))).build());
}
}Tables.toSentences and Markdown.toSpeech are your own helpers. The table rewrite parses the header row and emits one sentence per row ("Basic: price $10, data 5 GB."), which a screen reader reads in order and a listener can follow. The speech conversion strips emphasis markers, turns list bullets into sentences, expands symbols that text-to-speech engines read badly ("&" to "and", "/" to "or" where appropriate) and removes URLs in favour of their link text.
Some violations should be measured, not repaired. Rewriting "the red button" automatically needs knowledge of the real button label, so the callback only counts it; the counts tell you which instruction rule the model ignores, and you fix the instruction or the tool output that prompted it.
Streaming changes the picture. If your client shows partial chunks as they arrive, the user has already heard the table before the final response is repaired. For screen reader and voice profiles, either disable partial display and show a status message while the agent works, or buffer per sentence in the client and run the same conversion there. Buffering costs perceived latency, so apply it per profile.
Images in and out
Multimodal agents can be an accessibility tool in their own right. A blind user who photographs a damaged parcel or a bill can ask the agent what it shows; the instruction's last rule asks for descriptions focused on the user's question rather than a generic caption. Be honest about limits: tell users when the image is unclear, and never state a figure such as an amount due unless it is legible in the image.
In the other direction, anything non-text the agent returns needs a text alternative under 1.1.1. If a tool generates a chart as an artifact, have the tool also return a one-sentence summary of what the chart shows, and send both parts together, for example Content.fromParts(Part.fromText(summary), chartPart). The client uses the summary as the image's alternative text and as the spoken answer in voice mode.
The client: live regions, focus and timing
The client is where many accessibility failures happen, and ADK does not control it. Three rules cover most of the risk. First, announce completed units, not tokens: appending each streamed token to a live region makes some screen readers read fragments or restart constantly. Append completed sentences or the whole answer to a region with role="log", and use a separate role="status" element for "working" and "ready" messages (criterion 4.1.3). Second, keep focus where the user is: do not move focus to each new message; let users navigate the transcript with their screen reader's own commands. Third, everything works from the keyboard, the input has a real label, and timeouts warn before they expire and can be extended.
<!-- Completed sentences are appended here; a screen reader announces each addition once. -->
<div id="transcript" role="log" aria-live="polite" aria-relevant="additions"></div>
<p id="status" role="status"></p> <!-- "Agent is working", "Answer ready" -->
<form id="ask">
<label for="msg">Your question</label>
<textarea id="msg" required></textarea>
<button type="submit">Send</button>
</form>Honour the reducedMotion flag and the operating system's prefers-reduced-motion media query by turning off animated typing effects. For voice channels, keep turns short and offer to repeat; a listener cannot scroll back.
Worked example: a plan comparison
A screen reader user with {"screen_reader": true} asks which phone plan is cheaper for 15 GB of data. A tool returns plan data and the model, despite the instruction, answers with a markdown table followed by "Plus is the better fit". Rendered visually that is fine. Read aloud by a screen reader, a markdown table rendered as HTML is navigable but slow, and if the client renders raw text the user hears "vertical bar Plan vertical bar Price". The callback detects the table rows, records a table violation, and replaces the table with "Basic: price $10, data 5 GB. Plus: price $15, data 20 GB." The model's own conclusion follows unchanged.
Suppose a week of metrics shows the table rule firing on 9 percent of billing answers and almost never elsewhere. The cause is the billing tool, which returns rows; adding a summary field to its result, a sentence comparing the plans, drops the rate because the model now has prose to quote. The repair stays in place as the safety net.
Testing
Test each layer separately. Unit-test the callback with constructed contexts and responses, one test per rule and profile combination:
@Test
void tableIsRewrittenForScreenReaderUsers() {
CallbackContext ctx = contextWithState(Map.of(A11yProfile.STATE_KEY,
Map.of("screen_reader", true)));
LlmResponse in = textResponse("""
| Plan | Price | Data |
|------|-------|------|
| Basic | $10 | 5 GB |
| Plus | $15 | 20 GB |
""");
String out = textOf(A11yCallback.apply(ctx, in).blockingGet());
assertFalse(out.contains("|"));
assertTrue(out.contains("Basic: price $10, data 5 GB."));
}Add evaluation cases that run the whole agent with different profiles and assert properties of the final text: no pipes for screen reader profiles, no markdown at all for voice, no banned link phrases for anyone. Run an automated checker such as axe-core against the client to catch missing labels and roles, then test manually with real assistive technology: NVDA or JAWS on Windows, VoiceOver on macOS and iOS, TalkBack on Android. Best of all, pay disabled users to test it.
Failure modes
- Instruction-only accessibility. Works in demos, fails on the tool result nobody tested. Enforce format in code.
- Token-level live regions. The screen reader stutters or goes silent. Announce sentences or whole answers.
- Over-eager rewriting. A regex that converts every pipe character mangles code samples. Skip fenced code blocks in every rewrite.
- Profiles that leak. Storing the profile under a session key instead of
user:makes users repeat it every session. - Inferred disability. Guessing from behaviour is unreliable and intrusive. Let users choose, and default everyone to the accessible baseline.
- Fixed timeouts. Confirmation flows that expire in 30 seconds exclude slower users and fail 2.2.1.
What to do next
- Add the accessible-by-default rules to your agent instruction and review tool outputs for raw codes and tables.
- Define a minimal
user:profile with only output-changing settings and seed it from user preferences. - Implement an
afterModelCallbackthat records violations and repairs tables and markdown for affected profiles; see the callback architecture and guardrail ordering for how callbacks compose. - Make chart and image tools return a text summary alongside the image, following how ADK handles multimodal Parts.
- Fix the client: role=log region, separate status element, labelled input, keyboard operation, adjustable timeouts.
- Write unit tests per rule and profile and agent-level evaluation cases, then test with NVDA, VoiceOver and real users.
- Revisit how you write function tools so results carry human-readable messages.