Vertex AI Search, which Google's ADK documentation now also calls Agent Search, is Google Cloud's managed retrieval service: you load documents into a data store, it chunks, indexes and ranks them, and Gemini can ground its answers in what it retrieves. ADK for Java exposes it through one class, VertexAiSearchTool. Adding it to an agent takes four lines. Using it well takes an understanding of what those lines actually do, because the tool behaves very differently from the function tools you write yourself.
This article reads the tool from its source, explains the rules it enforces, shows how to combine it with other tools despite a documented one-built-in-tool-per-agent limit, covers per-tenant filtering, shows when to call the Search API yourself instead, and walks through an HR policy assistant from data store to answer. For the general theory of retrieval-augmented agents see RAG with ADK Java.
What the tool actually does
A FunctionTool is code ADK runs when the model emits a function call. VertexAiSearchTool is not. In the ADK Java source it extends BaseTool under the name vertex_ai_search and overrides processLlmRequest: before each model call it builds a VertexAISearch configuration from its fields, wraps it in a Retrieval tool, and appends that to the tools in the request's GenerateContentConfig. The model provider decides whether to search, runs the search against your data store, and writes the answer.
Three consequences follow. Your Java code never sees the raw search results, only what the model reports back in grounding metadata. Callbacks such as beforeToolCallback do not fire for the search, because no tool call passes through ADK. And everything about the search, which store, which filter, how many results, is fixed in the request configuration rather than chosen by the model at call time.
Preparing a data store
Before any Java runs you need a data store. In the Google Cloud console, create one under AI Applications, choose a source such as Cloud Storage, BigQuery or a website, and import documents. Unstructured documents such as PDFs can carry a metadata record with fields like tenant or department; mark the fields you will filter on as indexable in the schema, because filters on unindexed fields fail. Indexing takes time after an import, so build your evaluation only after documents report as indexed.
The tool needs the data store's full resource name, which has the form projects/<PROJECT>/locations/<LOCATION>/collections/default_collection/dataStores/<DATA_STORE_ID>. Many stores live in the global location; regional stores use their own location. If you have created a search app, also called an engine, over one or more stores, you can point the tool at the engine instead. The runtime needs a Google Cloud project and credentials rather than a Gemini API key: the ADK docs list GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION and, for Java, GOOGLE_APPLICATION_CREDENTIALS, plus a flag selecting the Vertex AI backend whose exact name you should copy from the current ADK setup guide. The broader platform setup is in Vertex AI on Google Cloud.
Wiring it into an agent
With a data store in place, the agent itself is short.
import com.google.adk.agents.LlmAgent;
import com.google.adk.tools.VertexAiSearchTool;
String dataStore = "projects/acme-hr/locations/global/collections/default_collection"
+ "/dataStores/hr-policies";
LlmAgent policyAgent = LlmAgent.builder()
.name("hr_policy_search")
.model("gemini-2.5-flash")
.description("Answers HR policy questions from the indexed policy library.")
.instruction("""
Answer only from the policy documents you retrieve.
Name the policy document you relied on.
If nothing relevant is retrieved, say you could not find it in the policies.""")
.tools(VertexAiSearchTool.builder()
.dataStoreId(dataStore)
.maxResults(5)
.build())
.build();The instruction does real work here. The model chooses whether to retrieve, so tell it that policy questions must be answered from retrieved documents, and give it a fixed way to say nothing was found. maxResults caps how many results feed the answer; fewer results mean shorter prompts and less irrelevant context, more mean better recall for broad questions. Start small and tune against an evaluation set.
The rules it enforces, and when they fail
The builder and the request hook enforce rules that are easy to trip over, and they fail at different times.
| Rule | Where it fails | What to do |
|---|---|---|
Exactly one of dataStoreId and searchEngineId | At build(), with an exception | Pick one; use an engine to search several stores |
dataStoreSpecs requires searchEngineId | At build() | Per-store specs only make sense for an engine |
| Model name must be a Gemini model | At request time, as an error from processLlmRequest | Test the agent end to end, not just its construction |
| Built-in tools stand alone in an agent | Documented ADK limitation for Java | Wrap the search agent with AgentTool |
The model check deserves emphasis. It inspects the model name and rejects anything that does not start with gemini or contain /gemini. Because the check happens when the request is prepared, an agent configured with another provider's model constructs fine and fails on its first turn. A unit test that only builds the agent will not catch it.
Combining search with other tools
ADK's documentation states that certain built-in tools, including Google Search, code execution and Agent Search, must be used by themselves in a single agent, and that this applies to ADK Java. Built-in tools also cannot be used inside sub-agents in Java. Real assistants need search alongside function tools, so the documented pattern is to put the search tool in its own agent and expose that agent to a parent as a tool.
import com.google.adk.tools.AgentTool;
import com.google.adk.tools.FunctionTool;
// The search agent holds the built-in tool alone; the parent holds everything else.
LlmAgent root = LlmAgent.builder()
.name("hr_assistant")
.model("gemini-2.5-flash")
.instruction("""
For policy questions call hr_policy_search with a focused question.
For leave balances call getLeaveBalance. Never guess a policy.""")
.tools(
AgentTool.create(policyAgent),
FunctionTool.create(LeaveTools.class, "getLeaveBalance"))
.build();
// ADK also ships a ready-made wrapper; note that it takes a BaseLlm, not a model name:
// VertexAiSearchAgentTool.create(model, VertexAiSearchTool.builder().dataStoreId(dataStore).build())The parent model now calls hr_policy_search like any function, passing a question, and receives the child agent's text answer. That has a cost: two model calls instead of one, and the parent sees a summary rather than the retrieved passages. If you need citations at the parent level, capture the child's grounding metadata in a callback, store it in session state, and have the parent or your UI read it from there. Callback mechanics are covered in ADK Java callbacks.
Scoping results per tenant
Filters use the Vertex AI Search filter syntax over indexable fields, for example department: ANY("finance") or a combination with AND. This is the right place for access control: a filter in the request is enforced by the search service, while an instruction asking the model to ignore other tenants' documents is a suggestion.
The catch is that the filter is a field on the tool, fixed at build time. ADK Python documents a subclassing hook for building the configuration dynamically; in Java the tool is an immutable value built once. There are two practical patterns. You can build one search agent per tenant and cache it, which works for a modest number of tenants. Or, when you need per-request filters, full control over ranking options, or access to raw results, call the Search API yourself from a function tool.
import com.google.adk.tools.Annotations.Schema;
import com.google.adk.tools.ToolContext;
import com.google.cloud.discoveryengine.v1.SearchRequest;
import com.google.cloud.discoveryengine.v1.SearchResponse.SearchResult;
import com.google.cloud.discoveryengine.v1.SearchServiceClient;
public final class PolicySearchTool {
// Copy the full serving config resource name from the console for your app or data store.
private static final String SERVING_CONFIG = System.getenv("POLICY_SERVING_CONFIG");
private static final SearchServiceClient CLIENT = create();
@Schema(name = "searchPolicies",
description = "Search the caller's company HR policies. Returns titles, links and ids.")
public static Map<String, Object> searchPolicies(
@Schema(name = "query", description = "A focused policy question") String query,
ToolContext ctx) {
// Tenant comes from trusted session state, never from the model's arguments.
String tenant = (String) ctx.state().get("tenant_id");
if (tenant == null || !tenant.matches("[a-z0-9-]{1,40}")) {
return Map.of("error", "no tenant bound to this session");
}
SearchRequest req = SearchRequest.newBuilder()
.setServingConfig(SERVING_CONFIG)
.setQuery(query)
.setPageSize(5)
.setFilter("tenant: ANY(\"" + tenant + "\")") // field must be indexed as filterable;
// add AND country: ANY(...) the same way
.build();
List<Map<String, String>> hits = new ArrayList<>();
for (SearchResult r : CLIENT.search(req).getPage().getValues()) {
var fields = r.getDocument().getDerivedStructData().getFieldsMap();
hits.add(Map.of(
"id", r.getDocument().getId(),
"title", fields.containsKey("title") ? fields.get("title").getStringValue() : "",
"link", fields.containsKey("link") ? fields.get("link").getStringValue() : ""));
}
return Map.of("results", hits);
}
// Regional (non-global) data stores may need a regional client endpoint set through
// SearchServiceSettings; check the client-library docs for your location.
private static SearchServiceClient create() {
try { return SearchServiceClient.create(); }
catch (IOException e) { throw new UncheckedIOException(e); }
}
}This is an ordinary FunctionTool, so callbacks fire, results are visible to your code, the tenant comes from trusted session state rather than from model-supplied arguments, and the tool can sit beside other tools in one agent. What you give up is built-in grounding metadata: the model sees results as a function response, so you assign ids and verify citations yourself, as described in grounding enforcement for ADK Java. The field names inside derived data depend on your data store type, so log one result before you rely on them.
Reading what was retrieved
With the built-in tool, citations arrive on the event. Event.groundingMetadata() returns an optional GroundingMetadata whose groundingChunks list retrieved documents, each with a retrieved context carrying title, URI and text; whose groundingSupports map segments of the answer to chunk indices; and whose retrievalQueries record the queries the model actually ran against your store. Log the retrieval queries during development: a model that searches for the wrong thing explains most poor answers, and no amount of index tuning fixes it.
Worked example: an HR policy assistant
Acme loads 340 HR policy PDFs into a data store, each with metadata fields country and tenant. An employee in Germany asks how much parental leave they can take. With the composed agent, the parent routes the question to hr_policy_search. Gemini issues a retrieval query such as "parental leave entitlement Germany", receives five results, and answers from the German leave policy, naming it.
In testing, the team sees a wrong answer for an employee in Austria: the model cited the German policy. The logged retrieval query lacked the country, and the German document ranked first. Prompt changes helped only intermittently. The fix was structural: employees' country is in session state, so the team switched to the function-tool pattern with a filter on both tenant and country. Wrong-country citations stopped, and the evaluation set gained ten cross-country questions to keep it that way.
Failure modes
- Error on first turn only. A non-Gemini model name passes construction and fails at request time.
- Build-time exception. Both or neither of data store and engine ids set, or store specs without an engine.
- Permission denied at search time. The runtime identity cannot read the data store. Discovery Engine Viewer is the usual grant; prove it with a deliberately unprivileged identity in staging.
- Empty or irrelevant results. Documents not yet indexed, a filter on an unindexed field, or the wrong location in the resource name.
- Tool silently not used. The model answered from memory. Tighten the instruction and track the share of answers with grounding metadata.
- Cross-tenant leakage. Tenant isolation was requested in the prompt instead of enforced in a filter.
Operating it and choosing a pattern
Operate it like any dependency. Track grounded-answer rate, retrieval queries per turn, latency of the search and model calls separately, and abstentions. Keep an evaluation set with answerable, unanswerable and cross-tenant questions, and rerun it after every import. Search and model calls are billed and rate limited separately, so check current quotas and pricing in your project rather than assuming limits. Choose the built-in tool for the shortest path with managed citations; choose the function-tool pattern when you need dynamic filters, observability of raw results or composition without an extra agent. Broader design guidance across languages is in ADK retrieval and grounding.
What to do next
- Create a data store, import a representative sample of documents, and mark the fields you will filter on as indexable.
- Build the four-line agent with
VertexAiSearchTooland a strict instruction, and run it end to end so the Gemini model check executes. - Log
retrievalQueriesand grounding chunks for every test turn. - If the agent needs other tools, move search into its own agent and expose it with
AgentTool. - Enforce tenant and permission scoping with filters, using per-tenant agents or a function tool that calls the Search API.
- Write an evaluation set including unanswerable and cross-tenant questions, and rerun it after every import.
- Grant the runtime identity read access only, and test the denied path deliberately.