A Java agent that answers from your documents needs somewhere to keep embeddings and a way to search them by meaning. Spring AI's answer is the VectorStore interface: one abstraction over pgvector, Redis, Elasticsearch, OpenSearch, Qdrant, Milvus, Weaviate, Pinecone, MongoDB Atlas, Cassandra, Neo4j, Oracle and others, with a portable metadata filter language, an ETL pipeline for loading documents, and advisors that plug retrieval into a ChatClient. Used well, it lets you switch databases by swapping a starter. Used carelessly, it hides the choices (distance metric, dimensions, chunking, filters) that decide whether retrieval works.

This article explains the abstraction from first principles, wires a pgvector store, builds an ingestion pipeline with safe re-ingestion, shows tenant-scoped search and the two RAG advisors, then covers observability, failure modes and trade-offs. Names and defaults were checked against the Spring AI 2.0.1 reference documentation; Spring AI moves quickly, so recheck them against the version in your BOM. For the ADK memory side of the same database, see Semantic Memory + Vector Stores for ADK Java Agents.

The abstraction: VectorStore, Document and the embedding model

A vector store does three things: it stores text with metadata and an embedding, it deletes entries, and it returns the entries whose embeddings are closest to a query's embedding, optionally restricted by a metadata filter. Spring AI splits this into two interfaces. VectorStoreRetriever has the read side, and VectorStore extends it together with DocumentWriter, which is a Consumer<List<Document>>:

public interface VectorStore extends DocumentWriter, VectorStoreRetriever {
    default String getName() { return this.getClass().getSimpleName(); }
    void add(List<Document> documents);
    void delete(List<String> idList);
    void delete(Filter.Expression filterExpression);
    default void delete(String filterExpression) { ... }
    default <T> Optional<T> getNativeClient() { return Optional.empty(); }
}

public interface VectorStoreRetriever {
    List<Document> similaritySearch(SearchRequest request);
    default List<Document> similaritySearch(String query) { ... }
}

The unit of storage is Document: an id, text (getText()) or media, a metadata map, and after a search a score (getScore()). You never pass vectors yourself. The store holds an EmbeddingModel and embeds text on add and the query string on similaritySearch. That is convenient, and it is also the first trap: ingestion and query must use the same model, and changing the model means re-embedding everything. getNativeClient() is the escape hatch to the underlying client when you need a feature the abstraction does not expose.

Spring AI vector store: ingestion path (left) and query path (right)DocumentReaderTika, PDF, textTokenTextSplitterchunks + metadataBatchingStrategytoken-sized batchesVectorStoreadd / delete / similaritySearchadd(docs)EmbeddingModeltext to vectorembedPostgreSQL + pgvectorvector_store table, HNSWChatClientuser questionRAG advisorQuestionAnswer or RetrievalSearchRequestfilter expressiontenant == 'acme'Micrometerdb.vector.client.operationobservationsEvery store implements the same interface; the starter you add decides which database sits underneath.
Ingestion and query paths through the Spring AI VectorStore abstraction.

Wiring a pgvector store

Add spring-ai-starter-vector-store-pgvector (version from the spring-ai-bom), an embedding model starter, and a datasource. The store's properties live under spring.ai.vectorstore.pgvector:

spring:
  datasource:
    url: jdbc:postgresql://localhost:5432/kb
    username: kb_app
    password: ${KB_DB_PASSWORD}
  ai:
    vectorstore:
      pgvector:
        index-type: HNSW            # default; NONE = exact search, IVFFLAT also available
        distance-type: COSINE_DISTANCE
        dimensions: 1536            # must match the embedding model's output
        initialize-schema: false    # default false; create the table through migrations
        schema-name: public
        table-name: vector_store
        max-document-batch-size: 10000

Schema initialisation is off by default (earlier versions created it automatically), so own the table in Flyway or Liquibase. The documented layout is:

CREATE EXTENSION IF NOT EXISTS vector;
CREATE EXTENSION IF NOT EXISTS hstore;
CREATE EXTENSION IF NOT EXISTS "uuid-ossp";

CREATE TABLE IF NOT EXISTS vector_store (
    id uuid DEFAULT uuid_generate_v4() PRIMARY KEY,
    content text,
    metadata json,
    embedding vector(1536)      -- set to your model's dimension
);

CREATE INDEX ON vector_store USING HNSW (embedding vector_cosine_ops);

Three choices are fixed by this table. The dimension must equal the model's output, and pgvector's HNSW index supports at most 2,000 dimensions for the vector type, so a 3,072-dimension model needs a smaller output setting (anything else means a custom schema you must verify against the store). The operator class must match the distance type. And metadata is stored as JSON, which is what filters translate into. If you prefer code to properties, PgVectorStore.builder(jdbcTemplate, embeddingModel) exposes the same settings (dimensions, distanceType, indexType, initializeSchema, schemaName, vectorTableName, maxDocumentBatchSize).

Ingestion: readers, splitters, metadata and re-ingestion

Spring AI's ingestion pipeline is three functional interfaces: DocumentReader (a Supplier), DocumentTransformer (a Function) and DocumentWriter (a Consumer, which every vector store is). The canonical pipeline is one line, vectorStore.accept(splitter.apply(reader.get())). Production ingestion needs two more things: metadata you can filter on, and a way to replace a document's chunks when it changes, so stale chunks do not linger next to new ones.

@Service
public class HandbookIngestor {
    private final VectorStore store;
    private final TokenTextSplitter splitter = TokenTextSplitter.builder()
            .withChunkSize(500)            // tokens; default 800
            .withMinChunkSizeChars(200)
            .withKeepSeparator(true)
            .build();

    public HandbookIngestor(VectorStore store) { this.store = store; }

    public int ingest(Resource file, String tenant, String sourceId, String version) {
        TikaDocumentReader reader = new TikaDocumentReader(file);
        List<Document> chunks = splitter.apply(reader.get()).stream()
                .map(d -> d.mutate()
                        .metadata("tenant", tenant)
                        .metadata("source", sourceId)
                        .metadata("version", version)
                        .build())
                .toList();
        FilterExpressionBuilder b = new FilterExpressionBuilder();
        // Replace, do not append: remove the previous chunks of this source first.
        store.delete(b.and(b.eq("tenant", tenant), b.eq("source", sourceId)).build());
        store.add(chunks);
        return chunks.size();
    }
}

Delete-then-add leaves a short window where the source has no chunks; if that matters, add the new version first and delete with version != 'new' afterwards. Batching happens inside add: the default TokenCountBatchingStrategy groups documents so each embedding request stays under a token budget (8,191 tokens by default, with a 10 percent reserve), and it throws if a single document exceeds the limit. If your embedding model has a different input limit, declare your own BatchingStrategy bean; Spring Boot then uses it for every vector store.

Searching: SearchRequest and the filter language

Queries go through SearchRequest. The defaults matter: topK is 4 and similarityThreshold is 0.0, which accepts everything. A query with no relevant documents therefore still returns four documents, and the model will try to use them. Set a threshold, but calibrate it on your data, because scores from different models and distance metrics are not comparable.

FilterExpressionBuilder b = new FilterExpressionBuilder();
SearchRequest request = SearchRequest.builder()
        .query("How many days of parental leave do contractors get?")
        .topK(6)
        .similarityThreshold(0.55)          // calibrated on a labelled query set
        .filterExpression(b.and(
                b.eq("tenant", tenantId),       // from the session, never from the user
                b.in("doc_type", "policy", "faq")).build())
        .build();
List<Document> hits = store.similaritySearch(request);
hits.forEach(d -> log.info("{} {} {}", d.getId(), d.getScore(), d.getMetadata().get("source")));

Filters can also be strings, such as genre == 'drama' && year >= 2020, and support comparison, IN/NIN, AND/OR/NOT and IS NULL/IS NOT NULL (the null checks are not implemented by every store). Each store translates the portable expression into its own query language; for pgvector that is a JSON path predicate. Prefer the builder for anything that contains user-influenced values: concatenating a string filter from input is the vector-store version of SQL injection, and a crafted value can widen a tenant filter.

Retrieval in the chat path: the RAG advisors

You can call similaritySearch yourself, but usually retrieval is wired into ChatClient through an advisor. QuestionAnswerAdvisor (artifact spring-ai-vector-store-advisor) runs a search for the user's text and adds the results to the prompt. RetrievalAugmentationAdvisor (artifact spring-ai-rag) is modular: a VectorStoreDocumentRetriever plus optional query transformers and a query augmenter. Both accept a filter per request, which is how you scope retrieval to a tenant:

Advisor rag = RetrievalAugmentationAdvisor.builder()
        .documentRetriever(VectorStoreDocumentRetriever.builder()
                .vectorStore(store)
                .similarityThreshold(0.55)
                .topK(6)
                .build())
        .queryAugmenter(ContextualQueryAugmenter.builder()
                .allowEmptyContext(false)   // default: refuse rather than answer with no context
                .build())
        .build();

String answer = chatClient.prompt()
        .advisors(rag)
        .advisors(a -> a.param(VectorStoreDocumentRetriever.FILTER_EXPRESSION,
                "tenant == '" + TenantIds.requireSafe(tenantId) + "'"))
        .user(question)
        .call()
        .content();

The request-time filter overrides a filter set on the retriever. Because this parameter is a string, validate the tenant id against a strict pattern (TenantIds.requireSafe stands for your own validator) before interpolating it.

In an ADK agent, expose retrieval as a function tool instead, following the idiom in ADK Java RAG Implementation. The model chooses when to search; the application, not the model, decides whose documents it can see, because the tenant comes from session state rather than from a tool parameter:

import com.google.adk.tools.Annotations.Schema;
import com.google.adk.tools.FunctionTool;
import com.google.adk.tools.ToolContext;     // ADK's ToolContext, not Spring AI's

public final class HandbookSearchTools {
    private final VectorStore store;
    public HandbookSearchTools(VectorStore store) { this.store = store; }

    @Schema(description = "Search the employee handbook. Returns passages with ids; cite them.")
    public Map<String, Object> searchHandbook(
            @Schema(name = "query", description = "A focused search query") String query,
            ToolContext toolContext) {
        Object tenant = toolContext.state().get("tenant_id");    // written by the app
        if (!(tenant instanceof String tenantId) || tenantId.isBlank()) {
            return Map.of("status", "error", "error_code", "NO_TENANT");
        }
        if (query == null || query.isBlank() || query.length() > 500) {
            return Map.of("status", "error", "error_code", "INVALID_QUERY");
        }
        FilterExpressionBuilder b = new FilterExpressionBuilder();
        List<Document> hits = store.similaritySearch(SearchRequest.builder()
                .query(query).topK(6).similarityThreshold(0.55)
                .filterExpression(b.eq("tenant", tenantId).build())
                .build());
        return Map.of("status", "success", "passages", hits.stream().map(d -> Map.of(
                "id", d.getId(),
                "source", String.valueOf(d.getMetadata().get("source")),
                "text", Objects.toString(d.getText(), ""))).toList());
    }
}

// LlmAgent.builder()...tools(FunctionTool.create(new HandbookSearchTools(store), "searchHandbook"))

The text is read with Objects.toString(d.getText(), "") because getText() is null for media documents and Map.of rejects nulls. Return only id, source and text: the function response is stored in the session and re-sent on later turns, so every extra field costs tokens on every turn. An empty passage list is a valid answer, and the agent's instruction should tell it to say it does not know rather than guess.

Observability

Every Spring AI vector store is instrumented with Micrometer. Calls produce db.vector.client.operation observations (Prometheus series such as db_vector_client_operation_seconds_count) with low-cardinality tags db.operation.name (add, delete, query), db.system (for example pg_vector) and spring.ai.kind. Traces carry high-cardinality details: top-k, threshold, filter, dimension count and the query text. Response documents are not exported unless you set spring.ai.vectorstore.observations.log-query-response=true, which is off by default for good reason: it writes document content, possibly personal data, into logs. Alert on query latency percentiles, on the share of queries returning zero hits after the threshold, and on add failures from the embedding provider.

Remember that a query's latency has two parts: the embedding call to the model provider and the database search. The vector store observation covers the whole operation, so compare it with the embedding model's own observations before blaming the index. A slow p99 caused by the provider calls for caching query embeddings or a closer region; a slow p99 caused by the database calls for EXPLAIN, index parameters or more memory. Track ingestion too: chunks added per source, embedding tokens spent, and the age of the newest version per source, so stale documents are visible before users find them.

Failure modes

  • Dimension mismatch. Changing the embedding model or its output size without recreating the column makes inserts fail; worse, a same-size different model inserts fine and silently returns nonsense. Store the model id in metadata and refuse to mix.
  • Threshold left at 0.0. Every query returns top-k results, relevant or not, and the model answers confidently from them.
  • Duplicate chunks. Re-running ingestion without deleting by source doubles results and skews ranking.
  • Leaky filters. A missing tenant clause, or one built by string concatenation from user input, returns other customers' documents.
  • Oversized documents. A chunk above the batching limit throws during add; split before writing, and decide whether the job skips or fails.
  • Index not used. An operator class that does not match the distance type makes PostgreSQL fall back to a sequential scan; check the plan with EXPLAIN.

Trade-offs

The abstraction buys portability, a common filter language, consistent observability and drop-in RAG advisors. It costs you the features it does not model: hybrid keyword plus vector search, per-query HNSW tuning such as hnsw.ef_search, quantised columns and partitioning. Reach for getNativeClient() or plain SQL for those, as pgvector at 100M Scale shows. As for which database: if you already run PostgreSQL, pgvector keeps vectors next to the rows they describe and inside one backup and access model; a dedicated vector database earns its keep at large scale or with demanding filtered search. Memory and Vector Store Options for Agentic Systems compares the options.

What to do next

  1. Pin your Spring AI version in the BOM and confirm each property and artifact named here against that version's reference docs.
  2. Create the vector_store table through a migration, with the dimension and operator class matching your embedding model and distance type.
  3. Build ingestion with tenant, source, version and model metadata, and replace chunks per source rather than appending.
  4. Label 50 to 100 real questions with the documents that answer them, measure recall at k, and pick topK and the similarity threshold from that data.
  5. Build every tenant filter with FilterExpressionBuilder or a validated id, and add a test that a second tenant's documents never appear.
  6. Scrape db.vector.client.operation metrics and alert on latency and zero-hit rates; keep log-query-response off in production.
Key takeaway: Spring AI's VectorStore puts one interface over many vector databases: add and delete Documents, and run similarity searches with a portable metadata filter, while the store embeds text with its EmbeddingModel. Fix the dimension, distance type and schema deliberately, replace chunks per source on re-ingestion, set and calibrate a similarity threshold, build tenant filters with the builder, and watch the db.vector.client.operation metrics. Drop to the native client for hybrid search and index tuning.