Vertex AI Search is Google Cloud's managed retrieval service. You point it at documents, website content or structured records. It parses, chunks, embeds and indexes them, and it serves ranked results or generated answers with citations through one API. For many retrieval-augmented generation (RAG) projects it replaces a hand-built stack of parser, chunker, embedding model, vector database, keyword index and reranker.
A note on names before anything else, because it affects every search you run on the docs. As of October 2026, Google's documentation says Vertex AI Search is being renamed Agent Search. It lists former names including Vertex AI Search, AI Applications, Agent Builder, Vertex AI Search and Conversation, Enterprise Search and Generative AI App Builder. The API is the Discovery Engine API (discoveryengine.googleapis.com), and the docs say app and engine mean the same thing in the API. This article uses the original name because that is how most code and tutorials refer to it.
This article explains the resource model, ingestion, parsing and chunking, the search and answer methods, filters and boosts, the check grounding API, a worked example, failure modes and trade-offs. API details were checked against the Google Cloud docs on 2026-10-04. Pricing and edition details change often and are left out, so check the current pricing page.
The resource model
Four resources make up a deployment:
- Data store. This is where content is ingested and indexed. A data store holds one kind of content. The docs list custom data (structured and unstructured), public websites, media content and healthcare FHIR data. Unstructured documents are typically imported from Cloud Storage, and structured records from BigQuery or JSON.
- App (engine). The thing you query. An app attaches one or more data stores. That attachment is how you search a document store and a structured catalogue with one request.
- Serving config. A named configuration under the app that requests go through. Requests address it directly, for example
servingConfigs/default_search. - Location and collection. Resource names include a location and a collection, usually
globalanddefault_collection. The location chosen at creation determines the endpoint you call.
The full resource name of the serving config is what every query uses:
projects/PROJECT_ID/locations/global/collections/default_collection/
engines/APP_ID/servingConfigs/default_search
Ingestion, parsing and chunking
Retrieval quality is mostly decided at ingestion time. For unstructured documents, Vertex AI Search offers three parsers:
| Parser | What it does | Use it for |
|---|---|---|
| Digital (default) | Extracts machine-readable text and detects text blocks, but not tables, lists or headings | Clean, text-first documents |
| OCR | Recognises text in scanned and image-based PDFs; useNativeText merges native text with OCR output | Scans, faxes, image PDFs |
| Layout | Detects titles, headings, tables and lists in PDF, HTML, DOCX, PPTX, XLSX and XLSM files | Structured documents such as manuals, policies and spreadsheets |
Chunking splits documents into retrievable passages. It is configured on the data store's documentProcessingConfig, alongside the parser:
"documentProcessingConfig": {
"defaultParsingConfig": { "layoutParsingConfig": {} },
"chunkingConfig": {
"layoutBasedChunkingConfig": {
"chunkSize": 300,
"includeAncestorHeadings": true
}
}
}Per the docs, chunkSize accepts 100 to 500 tokens and defaults to 500. includeAncestorHeadings includes the ancestor headings with each chunk, so a passage that says only "the limit is 30 days" still carries "Expenses > International travel" with it. Parser choice can also be overridden per file type with parsingConfigOverrides.
The most important operational fact: the docs state that document chunking cannot be turned on or off after data store creation. Decide on chunking before the first import. If you get it wrong, plan to create a new data store and re-import, then switch the app over.
Choosing a chunk size is the usual trade-off. Smaller chunks are more precise and pack more distinct passages into an LLM prompt, but they lose context and split tables. Larger chunks keep context but dilute relevance. Policy and FAQ content with short self-contained answers suits 200 to 300 tokens with ancestor headings. Long technical explanations suit the 500 maximum.
The search method
The search method returns ranked results. A minimal REST call:
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://discoveryengine.googleapis.com/v1/projects/PROJECT_ID/locations/global/collections/default_collection/engines/APP_ID/servingConfigs/default_search:search" \
-d '{
"query": "international travel expense limit",
"pageSize": 10,
"userPseudoId": "u-7f3a91",
"filter": "doc_type: ANY(\"policy\") AND region: ANY(\"EU\", \"GLOBAL\")",
"contentSearchSpec": {
"searchResultMode": "CHUNKS",
"snippetSpec": { "returnSnippet": true }
}
}'The fields worth knowing:
searchResultModeselectsCHUNKSorDOCUMENTS. Use chunks when the results feed an LLM, and documents when a human sees a results page.snippetSpec,extractiveContentSpecandsummarySpeccontrol what each result carries: highlighted snippets, extracted passages, or a generated summary of the top results with optional citations.userPseudoIdis a pseudonymous, stable visitor ID of at most 128 characters. Never send an email address or account ID. The docs note that self-learning ranking depends on user event data, and a consistent pseudonymous ID is what ties those events together.queryExpansionSpecandspellCorrectionSpeccontrol query rewriting. Consider disabling expansion for exact-match workloads such as part numbers.pageSizeandoffsetpaginate.
In Python, the google-cloud-discoveryengine client exposes the same request:
from google.cloud import discoveryengine_v1 as discoveryengine
client = discoveryengine.SearchServiceClient()
serving_config = (
"projects/PROJECT_ID/locations/global/collections/default_collection/"
"engines/APP_ID/servingConfigs/default_search")
request = discoveryengine.SearchRequest(
serving_config=serving_config,
query="international travel expense limit",
page_size=10,
filter='doc_type: ANY("policy")',
user_pseudo_id="u-7f3a91",
)
for result in client.search(request): # the pager follows next pages
print(result.document.id)
Filters, boosts and tenant scoping
Filters restrict results to documents whose metadata matches an expression. A field can be filtered only if it is marked Indexable in the data store schema, so decide your filter fields before importing. The docs describe these functions: ANY for exact match against literals, EXISTS, CONTAINS for token matches, STARTS_WITH, and IN for numeric ranges. Expressions combine with AND and OR, and NOT or a leading - negates.
Boosts change ranking without excluding anything. Each condition uses filter syntax, and its boost ranges from -1 (strong demotion) to 1 (strong promotion):
"boostSpec": {
"conditionBoostSpecs": [
{ "condition": "doc_type: ANY(\"policy\")", "boost": 0.5 },
{ "condition": "status: ANY(\"archived\")", "boost": -0.7 }
]
}A filter is not access control unless your backend owns it. If a browser or an LLM agent builds the filter string, a user or an injected prompt can change it. For multi-tenant or permission-sensitive corpora, call the API from your backend. Derive the tenant and group filter from the authenticated identity, and AND it onto anything the client supplies. ADK Java with Vertex AI Search shows the same per-tenant scoping pattern inside an agent.
The answer method and check grounding
The answer method runs the whole RAG loop on Google's side. It understands the query, searches, generates an answer and attaches citations. It uses the same serving config with :answer:
POST .../engines/APP_ID/servingConfigs/default_search:answer
{
"query": { "text": "Can I expense a rail upgrade on a 4-hour trip?" },
"searchSpec": { "searchParams": {
"maxReturnResults": 8,
"filter": "doc_type: ANY(\"policy\")" } },
"answerGenerationSpec": {
"includeCitations": true,
"ignoreAdversarialQuery": true,
"ignoreNonAnswerSeekingQuery": true,
"ignoreLowRelevantContent": true,
"promptSpec": { "preamble": "Answer only from the policy text. Quote the limit." }
}
}maxReturnResults defaults to 10, with a maximum of 25. The three ignore* flags make the service decline adversarial queries, chit-chat and questions the content does not answer well, instead of producing an unsupported answer. Pass session for multi-turn conversations, and use modelSpec.modelVersion to pin the generation model rather than float with the default. The response carries answer.state, answer.answerText, answer.citations and answer.references. Always check the state before showing text.
Search plus your own model, or the answer method? The answer method is fastest to ship and keeps grounding, citation and query classification in one call. Calling search and then your own model, for example Gemini as described in the Gemini API article, gives full control of the prompt, output format, tool use and model choice. Agents usually need that control.
If you take the second route, the check grounding API can score your model's answer against the passages you supplied. You send an answerCandidate of up to 4,096 tokens and up to 200 facts of up to 10,000 characters each to groundingConfigs/default_grounding_config:check. It returns a supportScore from 0 to 1, citedChunks, and per-claim citations. Claims are marked when they need no grounding. citationThreshold defaults to 0.6. Use the support score as a gate: below your calibrated threshold, regenerate or abstain.
Worked example: an engineering docs assistant
An engineering organisation wants an assistant over 40,000 internal documents: PDFs of design reviews, HTML runbooks and XLSX capacity sheets. Each document already has metadata for owning team, type and confidentiality.
- Data store. Create an unstructured data store with the layout parser, because tables in runbooks and sheets matter. Set chunking to 300 tokens with ancestor headings. Mark
team,doc_typeandconfidentialityas indexable. - Import. Load from Cloud Storage with metadata alongside each file. Schedule a re-import for changed files, and record the last successful import time as a metric.
- Backend. The agent never calls the API directly. A backend service maps the caller's groups to a filter such as
confidentiality: ANY("public", "internal"), plus a team filter for restricted teams, and ANDs it onto any filter the agent proposes. - Retrieval. The agent calls search in chunk mode with 8 results, then generates with its own model so that it can mix in tool calls.
- Gate. The answer and its chunks go to check grounding. Answers below a support score of 0.7 (chosen on a labelled set of 300 questions) are regenerated once, then replaced by "I could not find this in the docs" with the top three links.
- Evaluate. A golden set of 300 real questions with known source documents measures recall@8 at the chunk level and answer correctness weekly, and again after any change to chunking, filters or boosts.
The first evaluation run typically exposes ingestion problems rather than ranking problems: scanned PDFs with no text because the digital parser was used, tables flattened into prose, and headings missing from chunks. Each of these is fixed in the data store configuration, which is why the chunking decision has to be right before import.
Failure modes
Failure modes seen in practice:
- Locked chunking. Chunking was not enabled, or was set up wrongly, at creation. You cannot toggle it later, so plan a parallel data store and a cut-over.
- Filters silently ineffective. A field was not marked indexable, or the filter uses a different value spelling than the metadata. Test every filter against a known document.
- Client-built filters. Tenant isolation depends on a string the client controls. Build it in the backend from identity.
- Stale index. Imports fail quietly or website content is refreshed late, so answers cite superseded pages. Alert on import freshness and show the document's update date with citations.
- Ignoring the answer state. Code shows
answerTextwithout checkinganswer.stateor whether the query was skipped as adversarial or non-answer-seeking. - Personal data in
userPseudoId. Use a salted hash or a random per-user ID. - Wrong location. The location in the resource name, and the endpoint you call, must match where the data store and app were created. Copying a
globalexample for a regional deployment is a common cause of confusing errors.
Trade-offs
Managed versus self-built. Vertex AI Search removes the work of running parsers, embeddings, hybrid indexes and reranking, and it is strong on messy enterprise documents. You give up control of the embedding model, the fusion of keyword and vector scores, and the chunking algorithm beyond its parameters. If you need custom embeddings, an unusual similarity metric or tight latency in your own VPC, compare it with a self-built stack as described in hybrid search with BM25, dense retrieval and reranking.
Answer method versus search plus your model. The answer method is simpler but less flexible. Search plus your model costs an extra call but lets you control prompts, tools and output. Chunks versus documents. Use chunks for LLM context and documents for human results pages. Where it sits on Google Cloud. It complements, rather than replaces, the model hosting and pipelines covered in the Vertex AI architecture overview.
What to do next
- Check which name your console and docs use (Vertex AI Search, AI Applications or Agent Search) so your searches find current pages.
- Choose parser and chunking settings, and mark filter fields indexable, before the first import.
- Import a representative sample and inspect the chunks before loading the full corpus.
- Build a golden question set and measure chunk recall before tuning boosts or prompts.
- Call the API from a backend that derives filters from user identity.
- Choose between the answer method and search plus your own model, and pin the model version.
- If you generate answers yourself, gate them with check grounding at a calibrated threshold.
- Monitor import freshness, answer states and grounding scores, and re-run evaluation after every configuration change.