Retrieval-augmented generation sounds like one feature, but in an ADK Java agent it is two programs with very different lives. An ingestion pipeline runs offline, turning documents into searchable chunks and keeping them current as documents change. A retrieval tool runs inside every agent turn, turning the model's question into a handful of relevant passages under a strict latency budget. Most RAG problems people blame on the model are actually bugs in one of these two programs: chunks cut in the middle of a table, stale passages from a document that was edited last week, or a search that misses an exact product code because it only used vectors.

This article builds both halves on PostgreSQL with pgvector, using the ADK Java function-tool APIs. The ADK Java RAG overview explains why grounding matters, the grounding article shows how to verify the citations the model produces, and semantic memory on pgvector covers per-user memory rather than a shared knowledge base. Here the focus is the pipeline: what to store, how to keep it fresh, how to search it, and how to tell whether it works.

Advertisement

Build or buy first

ADK Java ships a managed option. VertexAiRagRetrieval, in the com.google.adk.tools.retrieval package, is a retrieval tool that queries Google's managed RAG service, which the ADK documentation now calls Knowledge Engine and which was previously known as RAG Engine. You give it RAG resources and an optional vector distance threshold, and the service handles chunking, embedding and storage. If your documents can live in that service and your tenancy model maps onto its corpora, start there; it removes most of what follows.

Build it yourself when you need data in your own database, access control that mirrors an existing permission model, keyword matching on identifiers, or a vendor-neutral pipeline. The rest of this article assumes that call.

The architecture

DIY RAG for an ADK Java agent: an offline ingestion path and an online retrieval toolIngestion (batch, idempotent)Query path (inside an agent turn)Source documentswiki, PDFs, tickets + ACL metadataStructure-aware chunkerheadings kept as breadcrumbsContent hash diffonly changed chunks re-embeddedEmbeddingClient (your seam)model id stored per rowPostgreSQL: kb_chunkpgvector HNSW + tsvector GINLlmAgent turnmodel decides to call searchDocsFunctionTool: searchDocstenant from session stateHybrid queryvector top 40 + keyword top 40RRF fusion, optional reranktop 6 chunks with idsupserttwo SQL scansfunction response: chunk ids + textThe model never touches the database: it can only ask the tool a question, and the tool applies the tenant filter itself.
Two programs share one table. Ingestion is batch and idempotent; retrieval is a function tool that the model calls during a turn and that enforces tenant scoping itself.

Keep the boundary sharp. Ingestion knows nothing about agents; the tool knows nothing about documents beyond rows. Their only contract is the schema, the embedding model id and the chunk id format, so each half is testable alone.

Advertisement

The schema

One table holds chunks. It carries both a vector column for semantic search and a generated tsvector column for keyword search, so both can be queried in one statement. The content hash makes ingestion incremental, and the embedding model identifier lets you migrate to a new model by writing new rows alongside old ones.

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE kb_chunk (
  tenant_id    text NOT NULL,
  chunk_id     text NOT NULL,             -- doc_id || ':' || ordinal, unique per tenant
  doc_id       text NOT NULL,
  doc_version  text NOT NULL,
  heading_path text NOT NULL,             -- "Billing > Refunds > Partial refunds"
  body         text NOT NULL,
  content_hash text NOT NULL,             -- sha256 of heading_path + body
  embed_model  text NOT NULL,
  embedding    vector(768) NOT NULL,      -- must equal the model's dimension
  tsv          tsvector GENERATED ALWAYS AS
                 (to_tsvector('english', heading_path || ' ' || body)) STORED,
  updated_at   timestamptz NOT NULL DEFAULT now(),
  PRIMARY KEY (tenant_id, chunk_id, embed_model)
);

CREATE INDEX kb_chunk_hnsw   ON kb_chunk USING hnsw (embedding vector_cosine_ops);
CREATE INDEX kb_chunk_tsv    ON kb_chunk USING gin (tsv);
CREATE INDEX kb_chunk_tenant ON kb_chunk (tenant_id, doc_id);

The cosine operator class on the HNSW index matches the <=> operator used in queries; mismatch them and the index is not used. A selective tenant filter combined with approximate search can return fewer rows than requested; the semantic memory article covers that filtered-search trap.

Chunking that respects structure

Chunk boundaries decide what the model can see. Fixed-size windows split tables, code samples and numbered procedures, and a chunk that says "step 4: confirm the amount" without the heading "Partial refunds" is nearly useless to retrieval. Split on the document's own structure first and on size second, and carry the heading path into every chunk.

public record Chunk(String id, String docId, String headingPath, String body) {}

public final class MarkdownChunker {
    private final int maxChars;     // e.g. 1800 chars, roughly 400-450 tokens of English
    private final int overlapChars; // e.g. 200, carried from the previous chunk

    public MarkdownChunker(int maxChars, int overlapChars) {
        this.maxChars = maxChars; this.overlapChars = overlapChars;
    }

    public List<Chunk> split(String docId, String markdown) {
        List<Chunk> out = new ArrayList<>();
        Deque<String> headings = new ArrayDeque<>();
        StringBuilder buf = new StringBuilder();
        for (String para : markdown.split("\n\\s*\n")) {
            Matcher m = Pattern.compile("^(#{1,6})\\s+(.*)").matcher(para.strip());
            if (m.find()) {                                   // new section: flush
                flush(out, docId, headings, buf);
                int level = m.group(1).length();
                while (headings.size() >= level) headings.removeLast();
                headings.addLast(m.group(2).strip());
                continue;
            }
            if (buf.length() + para.length() > maxChars && buf.length() > 0) {
                String tail = buf.substring(Math.max(0, buf.length() - overlapChars));
                flush(out, docId, headings, buf);
                buf.append(tail).append("\n\n");
            }
            buf.append(para.strip()).append("\n\n");
        }
        flush(out, docId, headings, buf);
        return out;
    }

    private void flush(List<Chunk> out, String docId, Deque<String> h, StringBuilder buf) {
        if (buf.toString().isBlank()) { buf.setLength(0); return; }
        out.add(new Chunk(docId + ":" + out.size(), docId,
                String.join(" > ", h), buf.toString().strip()));
        buf.setLength(0);
    }
}

The heading path is embedded and indexed with the body, so a query about partial refunds matches a chunk whose body never repeats those words. Overlap carries the end of a long section into the next chunk so a sentence split by the size limit still appears whole in one of them. Measure chunk length in tokens against your embedding model's input limit; depending on the API, longer text is rejected or silently truncated.

Incremental ingestion

Re-embedding a whole corpus every night is slow and expensive, and it hides bugs because everything always looks fresh. Hash each chunk, compare with what is stored, and embed only what changed.

public void ingest(String tenantId, String docId, String version, String markdown) throws SQLException {
    List<Chunk> chunks = chunker.split(docId, markdown);
    Map<String, String> existing = loadHashes(tenantId, docId, embedder.modelId()); // chunk_id -> hash

    List<Chunk> changed = new ArrayList<>();
    for (Chunk ch : chunks) {
        String h = sha256(ch.headingPath() + "\n" + ch.body());
        if (!h.equals(existing.get(ch.id()))) changed.add(ch);
    }
    List<float[]> vectors = changed.isEmpty() ? List.of()
            : embedder.embedDocuments(changed.stream()
                  .map(ch -> ch.headingPath() + "\n" + ch.body()).toList());

    try (Connection cx = db.getConnection()) {
        cx.setAutoCommit(false);
        upsert(cx, tenantId, version, changed, vectors);  // ON CONFLICT (tenant_id, chunk_id, embed_model)
        deleteBeyond(cx, tenantId, docId, chunks.size()); // doc got shorter: drop stale tail chunks
        cx.commit();
    }
}

Two details matter. Deleting chunks beyond the new count handles documents that got shorter; without it, the old tail stays searchable forever. And the whole document is replaced in one transaction, so a query never sees half of version 7 mixed with half of version 8. Deleted documents need an explicit tombstone from the source system, or a periodic reconciliation that lists source ids and removes rows whose document no longer exists. Ingestion that only ever adds is the most common cause of answers quoting retired policy.

Hybrid retrieval with reciprocal rank fusion

Vector search is good at meaning and bad at exact strings. A user asking about error code E-4021 or SKU 88-1173 needs keyword matching, because embeddings blur rare tokens. Run both searches and merge the rankings. Reciprocal rank fusion is the simplest robust merge: each result scores the sum of one over sixty plus its rank in each list, so items ranked well by both rise to the top and the raw scores, which are on incomparable scales, never need normalising.

WITH vec AS (
  SELECT chunk_id, row_number() OVER (ORDER BY embedding <=> ?::vector) AS r
  FROM kb_chunk
  WHERE tenant_id = ? AND embed_model = ?
  ORDER BY embedding <=> ?::vector
  LIMIT 40
), kw AS (
  SELECT chunk_id,
         row_number() OVER (ORDER BY ts_rank_cd(tsv, q) DESC) AS r
  FROM kb_chunk, websearch_to_tsquery('english', ?) AS q
  WHERE tenant_id = ? AND embed_model = ? AND tsv @@ q
  ORDER BY ts_rank_cd(tsv, q) DESC
  LIMIT 40
), fused AS (
  SELECT chunk_id,
         COALESCE(1.0 / (60 + vec.r), 0) + COALESCE(1.0 / (60 + kw.r), 0) AS rrf
  FROM vec FULL OUTER JOIN kw USING (chunk_id)
)
SELECT c.chunk_id, c.doc_id, c.heading_path, c.body, f.rrf
FROM fused f
JOIN kb_chunk c ON c.chunk_id = f.chunk_id AND c.tenant_id = ? AND c.embed_model = ?
ORDER BY f.rrf DESC
LIMIT ?;

Pull more candidates than you return so fusion has something to work with. A cross-encoder reranker, if you add one, goes after fusion; keep it behind an interface so you can measure whether it earns its latency. The constant 60 comes from the original RRF paper.

The retrieval tool

Expose retrieval to the model as an instance function tool, so the retriever and its database connection are injected and can be mocked in tests. The tenant comes from session state that your application writes when it creates the session; it is never a tool parameter, because anything the model can set, a prompt injection can set too.

public final class DocSearchTools {
    private final HybridRetriever retriever;   // wraps the SQL above, injected, mockable

    public DocSearchTools(HybridRetriever retriever) { this.retriever = retriever; }

    @Schema(description = "Search the company knowledge base. Use for questions about "
            + "policies, products or procedures. Returns passages with ids; cite the ids you use.")
    public Map<String, Object> searchDocs(
            @Schema(name = "query", description = "A focused search query in plain words") String query,
            ToolContext toolContext) {
        Object tenant = toolContext.state().get("tenant_id");     // set by the app, never by the model
        if (!(tenant instanceof String tenantId) || tenantId.isBlank()) {
            return Map.of("status", "error", "error_code", "NO_TENANT",
                    "message", "Search is unavailable for this session.");
        }
        if (query == null || query.isBlank() || query.length() > 500) {
            return Map.of("status", "error", "error_code", "INVALID_QUERY",
                    "message", "Send a short, specific search query.");
        }
        List<Hit> hits = retriever.search(tenantId, query, 6);
        if (hits.isEmpty()) {
            return Map.of("status", "success", "passages", List.of(),
                    "note", "No matching documents. Say so rather than guessing.");
        }
        return Map.of("status", "success", "passages", hits.stream().map(h -> Map.of(
                "id", h.chunkId(), "section", h.headingPath(), "text", h.body())).toList());
    }
}

LlmAgent agent = LlmAgent.builder()
    .name("kb_assistant")
    .model("<a model id your project can use>")
    .instruction("Answer only from passages returned by searchDocs. Search before answering any "
        + "policy or product question; search again with different words if results look off-topic. "
        + "Cite passage ids in square brackets. If nothing relevant is found, say you do not know.")
    .tools(FunctionTool.create(new DocSearchTools(retriever), "searchDocs"))
    .build();

Return only id, section and text per passage: the function response is stored in the session and re-sent on later turns. Compile with -parameters so ADK sees the parameter names and recognises toolContext by name, as the function tool guide explains.

Worked example

A support agent for tenant acme receives: "Can I refund only the shipping on order 5512?" The model calls searchDocs with the query "refund shipping only partial refund". The vector search ranks the chunk billing-faq:3 (heading path Billing > Refunds > Partial refunds) first and a returns-policy chunk second. The keyword search, matching "shipping" and "refund", ranks shipping-policy:2 first and billing-faq:3 third.

Fusion scores billing-faq:3 at 1/61 + 1/63, about 0.0323. shipping-policy:2 is not in the vector top 40, so it scores 1/61, about 0.0164. The tool returns six passages led by the partial-refund procedure. The model answers with the steps and cites [billing-faq:3], and the grounding callback checks that cited id was actually returned this turn. Had the tenant key been missing from state, the tool would have refused before running a single query.

Measure retrieval separately from answers

Answer quality mixes retrieval, prompt and model. Retrieval on its own is cheap to measure and is where most regressions start. Keep a golden set of fifty to two hundred real questions, each labelled with the chunk ids that answer it, and compute recall at k on every change to chunking, embeddings or the query.

// golden.jsonl: {"tenant":"acme","query":"can I refund part of an order","relevant":["billing-faq:3"]}
double recallAtK(List<GoldenCase> cases, HybridRetriever r, int k) {
    int found = 0, total = 0;
    for (GoldenCase g : cases) {
        Set<String> got = r.search(g.tenant(), g.query(), k).stream()
                .map(Hit::chunkId).collect(Collectors.toSet());
        for (String rel : g.relevant()) { total++; if (got.contains(rel)) found++; }
    }
    return total == 0 ? 0 : (double) found / total;
}

Chunk ids change on re-chunking, so label cases by document and heading and resolve ids at test time. Track recall at 6 for the final list and recall at 40 for each candidate list: if the candidates contain the answer but the final list does not, fix fusion or reranking; if neither does, fix chunking or the embedding model.

Failure modes

SymptomCauseFix
Answers quote a retired policyIngestion never deletesTombstones or periodic reconciliation
Exact codes and SKUs not foundVector-only searchHybrid search with RRF
Steps without contextFixed-size chunks, no heading pathStructure-aware chunker with breadcrumbs
Another tenant's document appearsTenant taken from a tool argumentTenant only from session state, filtered in SQL
Fewer results than requestedSelective filter after approximate searchSee the filtered-search trap; partition or raise candidates
Slow turnsEmbedding call and two scans on the hot pathBound with a timeout; cache query embeddings
Quality drops after model swapOld and new vectors mixedFilter by embed_model; backfill, then switch

Put a timeout on the tool, as described in tool timeout handling, and return an error map the model can relay rather than letting the turn hang on a slow database.

What to do next

  1. Decide build or buy: try VertexAiRagRetrieval with Knowledge Engine first if your data and tenancy fit it.
  2. Create the kb_chunk table with both indexes and record the embedding model id on every row.
  3. Replace fixed-size chunking with a structure-aware chunker that carries heading paths.
  4. Make ingestion incremental with content hashes, and handle shorter and deleted documents explicitly.
  5. Add keyword search and reciprocal rank fusion, and expose retrieval as an instance FunctionTool that reads tenant from state.
  6. Build a golden set and gate every pipeline change on recall at k before looking at answer quality.
Key takeaway: RAG in ADK Java is an offline ingestion pipeline and an online retrieval tool joined by one table. Chunk on document structure with heading paths, re-embed only what changed and delete what disappeared, search with vectors and keywords fused by reciprocal rank, and expose the search as a function tool that takes its tenant from session state, never from the model. Measure recall on a golden set before you judge answers.