An agent that can read the web answers questions its training data cannot: this week's release notes, a vendor's current pricing page, the status of an incident. ADK Java gives you three ways to do it, and they differ in who fetches the page, what the model actually sees and how much control you keep. This article compares them and then goes deep on the two that ship in the library: UrlContextTool, where Gemini fetches pages on its own side, and ComputerUseToolset, where the model drives a real browser through screenshots and an interface you implement.
The third option, a custom fetch tool built on Playwright with an egress proxy, has its own article: ADK Java + Web Browser, in depth. Here it is the baseline the other two are measured against. Every ADK and genai detail below was read from the google-adk 1.11.0 and google-genai 1.75.0 jars with javap. Browsing APIs move quickly, so pin your versions and re-check the signatures when you upgrade.
Three ways an agent reads the web
Start from the question the agent must answer, not from the most capable tool. The three approaches sit on a ladder of power and risk:
| Approach | Who fetches | What the model sees | Your control | Use it for |
|---|---|---|---|---|
UrlContextTool | Google, server side | Page content folded into its context; you never see the bytes | Almost none: no allowlist hook, no per-URL status in ADK | Public documents named by URL |
Custom fetch FunctionTool | Your service | Text you extracted, labelled and truncated | Full: egress, size, caching, logging | Pages you must police or post-process |
ComputerUseToolset | Your browser, driven by the model | Screenshots plus the current URL | Full over the browser, but the action space is wide | JavaScript UIs, pickers, multi-page navigation |
The rule that falls out of the table: use the least powerful option that answers the question. Reading a known public URL is a fetch problem. Navigating an interface that only reveals content after clicks is a computer-use problem. Anything that signs in or changes state on a website is a computer-use problem that needs human confirmation, and is usually better solved with the site's API.
UrlContextTool: let Gemini fetch
In 1.11.0 the class is tiny. Its processLlmRequest copies the request's GenerateContentConfig, appends a genai Tool carrying an empty UrlContext, and writes the config back. Nothing runs on your side: no HTTP client, no allowlist, no model-name check. URLs in the conversation reach Gemini, Gemini retrieves them, and the answer comes back as ordinary text. Because ADK does not check the model, an unsupported model fails at the API call, not when you build the agent.
LlmAgent reader = LlmAgent.builder()
.name("page_reader")
.model("gemini-2.5-flash")
.description("Reads public web pages by URL and answers questions about them.")
.instruction("""
Answer only from the pages at the URLs you are given.
Quote the sentence you relied on and name its URL.
If a page cannot be retrieved, say so. Never guess its content.
""")
.tools(UrlContextTool.INSTANCE)
.build();
LlmAgent root = LlmAgent.builder()
.name("assistant")
.model("gemini-2.5-flash")
.instruction("Call page_reader when a question needs a specific public page.")
.tools(AgentTool.create(reader), FunctionTool.create(tickets, "lookupTicket"))
.build();Why a separate agent? Gemini's documentation says Gemini 3 models can combine built-in tools such as URL context with function calling, and does not promise it for older models. Wrapping the built-in tool in its own agent behind AgentTool works on any model and gives the reader a narrow instruction of its own, which matters because page text is untrusted input.
What you give up is visibility. genai's Candidate has a urlContextMetadata() accessor with per-URL retrieval status, but ADK's LlmResponse in 1.11.0 has no field for it; it carries content, grounding metadata, finish reason, error code, usage and a few others. A callback therefore cannot tell a page that loaded from one that failed. The instruction above asks the model to report failures in text; if you need the status as data, call the genai client from a custom tool where the whole response is in your hands. The documentation also says URLs must be publicly reachable, with no localhost or private networks, and it lists per-request and per-URL limits that you should check before designing around large batches.
ComputerUseToolset: the model drives a browser
Computer use is a loop. The model sees a screenshot, proposes an action such as a click at a coordinate, your code performs it, and a fresh screenshot goes back. ADK packages the loop as ComputerUseToolset, built from an object implementing BaseComputer. What the 1.11.0 bytecode shows:
- Two constructors:
ComputerUseToolset(BaseComputer)andComputerUseToolset(BaseComputer, int[]). The array sets the coordinate space the model works in, and the default is 1000 by 1000. - The toolset reflects over
BaseComputer's methods and exposes one tool per action, skipping the lifecycle and query methodsinitialize,close,screenSize,currentStateandenvironment. - Coordinates arrive in the model's space; ADK rescales them to what your
screenSize()reports and clamps them to the screen before calling you. - Every action returns
Single<ComputerState>: PNG screenshot bytes and an optional URL. The toolset sends the screenshot back to the model asimage/png.
The actions are openWebBrowser, clickAt, hoverAt, typeTextAt(x, y, text, pressEnter, clearBeforeTyping), scrollDocument, scrollAt, goBack, goForward, search, navigate, keyCombination and dragAndDrop. The interface also declares wait, but the 1.11.0 exclusion list matches by name and contains wait, so it is not exposed. ADK ships the interface, not a browser.
Implementing BaseComputer with Playwright
Playwright for Java is not thread-safe: a Playwright instance and everything created from it must be used from one thread. RxJava makes that easy, because subscribeOn a single-thread scheduler confines every action without locks.
public final class PlaywrightComputer implements BaseComputer {
private static final int W = 1280, H = 800;
private final ExecutorService thread =
Executors.newSingleThreadExecutor(r -> new Thread(r, "browser"));
private final Scheduler scheduler = Schedulers.from(thread);
private final Set<String> allowedHosts;
private Playwright pw; private Browser browser; private BrowserContext ctx; private Page page;
public PlaywrightComputer(Set<String> allowedHosts) { this.allowedHosts = allowedHosts; }
private Single<ComputerState> onPage(Consumer<Page> action) {
return Single.fromCallable(() -> {
action.accept(page);
page.waitForLoadState(LoadState.DOMCONTENTLOADED);
return ComputerState.create(page.screenshot(), page.url());
}).subscribeOn(scheduler);
}
@Override public Completable initialize() {
return Completable.fromAction(() -> {
pw = Playwright.create();
browser = pw.chromium().launch();
ctx = browser.newContext(new Browser.NewContextOptions().setViewportSize(W, H));
ctx.route("**/*", route -> { // defence in depth, not the boundary
String host = URI.create(route.request().url()).getHost();
if (host != null && allowedHosts.contains(host)) route.resume(); else route.abort();
});
page = ctx.newPage();
}).subscribeOn(scheduler);
}
@Override public Single<int[]> screenSize() { return Single.just(new int[] {W, H}); }
@Override public Single<ComputerEnvironment> environment() {
return Single.just(ComputerEnvironment.ENVIRONMENT_BROWSER);
}
@Override public Single<ComputerState> currentState() { return onPage(p -> {}); }
@Override public Single<ComputerState> openWebBrowser() { return currentState(); }
@Override public Single<ComputerState> navigate(String url) { return onPage(p -> p.navigate(url)); }
@Override public Single<ComputerState> clickAt(int x, int y) { return onPage(p -> p.mouse().click(x, y)); }
@Override public Single<ComputerState> typeTextAt(int x, int y, String text,
Boolean pressEnter, Boolean clearBeforeTyping) {
return onPage(p -> {
p.mouse().click(x, y);
if (Boolean.TRUE.equals(clearBeforeTyping)) p.keyboard().press("ControlOrMeta+A");
p.keyboard().type(text);
if (Boolean.TRUE.equals(pressEnter)) p.keyboard().press("Enter");
});
}
@Override public Single<ComputerState> scrollDocument(String direction) {
return onPage(p -> p.mouse().wheel(0, "up".equalsIgnoreCase(direction) ? -H : H));
}
// hoverAt, scrollAt, wait, goBack, goForward, search, keyCombination and
// dragAndDrop follow the same onPage(...) pattern.
@Override public Completable close() {
return Completable.fromAction(() -> { ctx.close(); browser.close(); pw.close(); })
.subscribeOn(scheduler)
.doFinally(thread::shutdown);
}
}Four decisions in that class matter more than the rest. screenSize() must report the real viewport, or ADK's rescaling sends clicks to the wrong place. Waiting for the load state before the screenshot keeps the model from reasoning about a half-painted page; for single-page apps you may need to wait for a specific element instead. One computer instance belongs to one session, never shared across users, because the browser context holds cookies and storage. And the route filter is only a second line: exact host matching misses subdomains and anything the browser fetches outside Playwright's routing, so put a real egress proxy in front, as the browser-tool article describes. The direction strings are passed through from the model, so handle left and right as well as up and down in production code.
ComputerUseToolset browserTools =
new ComputerUseToolset(new PlaywrightComputer(Set.of("docs.example.com")));
LlmAgent operator = LlmAgent.builder()
.name("browser_operator")
.model(COMPUTER_USE_MODEL) // a computer-use capable Gemini model; check the current list
.instruction("Find the requested fact on docs.example.com, then stop and report it.")
.tools(browserTools)
.beforeToolCallback(browseGuard)
.build();
Guarding the loop: budgets, allowlists and confirmation
Gemini's computer-use documentation says a proposed action can carry a safety_decision; when its decision is require_confirmation, the client is expected to ask the end user before acting. The 1.11.0 computer-use classes contain no string constant for that field, so assume ADK does not handle it and do it yourself. A beforeToolCallback is the natural place: it sees the tool and its arguments before ADK executes anything, and returning a map short-circuits the call, with the map becoming the tool result the model reads.
Callbacks.BeforeToolCallback browseGuard = (invocation, tool, args, toolContext) -> {
State state = toolContext.state();
int steps = ((Number) state.getOrDefault("browse_steps", 0)).intValue() + 1;
state.put("browse_steps", steps);
if (steps > 25) {
return Maybe.just(Map.of("error", "Step budget exhausted. Report what you found and stop."));
}
if (args.get("safety_decision") instanceof Map<?, ?> d
&& "require_confirmation".equals(d.get("decision"))) {
return Maybe.just(Map.of("error", "This action needs user confirmation and was not performed."));
}
// Tool names derive from BaseComputer method names; log them once at startup.
if (tool.name().toLowerCase(Locale.ROOT).contains("navigate")) {
String host = URI.create(String.valueOf(args.get("url"))).getHost();
if (host == null || !ALLOWED_HOSTS.contains(host)) {
return Maybe.just(Map.of("error", "Navigation to " + host + " is not allowed."));
}
}
return Maybe.empty(); // run the action
};The step budget lives in session state, so it survives across model calls in a turn; reset it when a new user message arrives. The confirmation branch refuses rather than pauses; to genuinely ask, end the turn with a question to the user and resume when they answer. Refusing with a clear message is still better than silently performing an action the model itself flagged. For a fuller treatment of the callback ordering, see ADK Java callbacks.
Worked example: a release note behind a picker
A user asks whether version 4.2 of an SDK hosted on docs.example.com deprecated its synchronous client. The root agent sees a public question about a known site and calls page_reader with the release-notes URL. Gemini fetches the page, finds the deprecation paragraph and answers with a quotation and the URL: one model call on the reader, one on the root, a few seconds end to end.
Now suppose the release notes sit behind a JavaScript version picker, and the fetched HTML contains only the picker. The reader reports that the page did not contain release notes for 4.2, as its instruction demands. The root falls back to browser_operator, and the trace looks like this:
openWebBrowserreturns a blank page screenshot; the guard counts step 1.navigateto the release-notes URL; the host is allowlisted; the screenshot shows the picker.clickAton the picker, thenclickAton 4.2; the screenshot shows the 4.2 notes.scrollDocumentdown twice until the API changes section is visible.- The model answers, quoting the line it read on screen, and names the final URL from
ComputerState.
Six actions means six model round trips, each carrying a fresh screenshot, so expect tens of seconds rather than a few, and many more input tokens; read the real figures from usageMetadata on the events rather than estimating. That cost is why the operator is the fallback, not the default.
Failure modes
- Unsupported model.
UrlContextTooladds the tool whatever the model; a model without URL context fails at request time. Cover it with a startup smoke test. - Silent retrieval failure. ADK hides per-URL status, so a model with a loose instruction may answer from memory as if it had read the page. Require quotations and test with a URL that returns 404.
- Prompt injection. Page text can contain instructions. With fetching it can only distort an answer; with computer use it can steer clicks. Allowlist hosts, never give the browser credentials, and confirm any typing into forms.
- Coordinate drift. If
screenSize()disagrees with the actual viewport or device scale factor, every click lands off target and the model retries forever. Assert the two match at startup. - Runaway loops. The model can scroll or click in circles; the step budget and a turn timeout (see tool timeout handling) bound the damage.
- Leaked browsers. A Chromium per session is hundreds of megabytes; close the toolset when the session ends and cap concurrent operators per replica.
Trade-offs
UrlContextTool is the cheapest to build and run, and the least observable: no logs of what was fetched, no egress control, no status. A custom fetch tool costs a service to operate and buys control and audit. Computer use handles interfaces the other two cannot, at the price of latency, token cost and a much larger attack surface. Most production agents need the first or second; keep the third behind its own agent, its own budget and its own allowlist. Whichever you choose, treat page content as data and keep the policy layer described in ADK Java guardrails in front of anything the agent does with what it read.
What to do next
- List the questions your agent must answer from the web and mark each as known-URL, search-then-read or needs-interaction.
- Put
UrlContextToolin a dedicated reader agent behindAgentTooland add an instruction that forces quotations and failure reports. - Write a smoke test that runs the reader against a real page and a 404 page on your chosen model.
- If you need interaction, implement
BaseComputeron a single-thread scheduler and assert thatscreenSize()matches the viewport. - Add the
beforeToolCallbackguard with a step budget, host allowlist and confirmation handling before the operator sees real traffic. - Route the browser through an egress proxy and log every navigation with the session ID.