Retrieval-augmented generation sounds like one feature, but in an ADK Java agent it is two programs with very different lives. An ingestion pipeline runs offline, turning documents into searchable chunks and keeping them current as documents change. A retrieval tool runs inside every agent turn, turning the model's question into a handful of relevant passages under a strict latency budget. Most RAG problems people blame on the model are actually bugs in one of these two programs: chunks cut in the middle of a table, stale passages from a document that was edited last week, or a search that misses an exact product code because it only used vectors.
This article builds both halves on PostgreSQL with pgvector, using the ADK Java function-tool APIs. The ADK Java RAG overview explains why grounding matters, the grounding article shows how to verify the citations the model produces, and semantic memory on pgvector covers per-user memory rather than a shared knowledge base. Here the focus is the pipeline: what to store, how to keep it fresh, how to search it, and how to tell whether it works.
Build or buy first
ADK Java ships a managed option. VertexAiRagRetrieval, in the com.google.adk.tools.retrieval package, is a retrieval tool that queries Google's managed RAG service, which the ADK documentation now calls Knowledge Engine and which was previously known as RAG Engine. You give it RAG resources and an optional vector distance threshold, and the service handles chunking, embedding and storage. If your documents can live in that service and your tenancy model maps onto its corpora, start there; it removes most of what follows.
Build it yourself when you need data in your own database, access control that mirrors an existing permission model, keyword matching on identifiers, or a vendor-neutral pipeline. The rest of this article assumes that call.
The architecture
Keep the boundary sharp. Ingestion knows nothing about agents; the tool knows nothing about documents beyond rows. Their only contract is the schema, the embedding model id and the chunk id format, so each half is testable alone.
The schema
One table holds chunks. It carries both a vector column for semantic search and a generated tsvector column for keyword search, so both can be queried in one statement. The content hash makes ingestion incremental, and the embedding model identifier lets you migrate to a new model by writing new rows alongside old ones.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE kb_chunk (
tenant_id text NOT NULL,
chunk_id text NOT NULL, -- doc_id || ':' || ordinal, unique per tenant
doc_id text NOT NULL,
doc_version text NOT NULL,
heading_path text NOT NULL, -- "Billing > Refunds > Partial refunds"
body text NOT NULL,
content_hash text NOT NULL, -- sha256 of heading_path + body
embed_model text NOT NULL,
embedding vector(768) NOT NULL, -- must equal the model's dimension
tsv tsvector GENERATED ALWAYS AS
(to_tsvector('english', heading_path || ' ' || body)) STORED,
updated_at timestamptz NOT NULL DEFAULT now(),
PRIMARY KEY (tenant_id, chunk_id, embed_model)
);
CREATE INDEX kb_chunk_hnsw ON kb_chunk USING hnsw (embedding vector_cosine_ops);
CREATE INDEX kb_chunk_tsv ON kb_chunk USING gin (tsv);
CREATE INDEX kb_chunk_tenant ON kb_chunk (tenant_id, doc_id);The cosine operator class on the HNSW index matches the <=> operator used in queries; mismatch them and the index is not used. A selective tenant filter combined with approximate search can return fewer rows than requested; the semantic memory article covers that filtered-search trap.
Chunking that respects structure
Chunk boundaries decide what the model can see. Fixed-size windows split tables, code samples and numbered procedures, and a chunk that says "step 4: confirm the amount" without the heading "Partial refunds" is nearly useless to retrieval. Split on the document's own structure first and on size second, and carry the heading path into every chunk.
public record Chunk(String id, String docId, String headingPath, String body) {}
public final class MarkdownChunker {
private final int maxChars; // e.g. 1800 chars, roughly 400-450 tokens of English
private final int overlapChars; // e.g. 200, carried from the previous chunk
public MarkdownChunker(int maxChars, int overlapChars) {
this.maxChars = maxChars; this.overlapChars = overlapChars;
}
public List<Chunk> split(String docId, String markdown) {
List<Chunk> out = new ArrayList<>();
Deque<String> headings = new ArrayDeque<>();
StringBuilder buf = new StringBuilder();
for (String para : markdown.split("\n\\s*\n")) {
Matcher m = Pattern.compile("^(#{1,6})\\s+(.*)").matcher(para.strip());
if (m.find()) { // new section: flush
flush(out, docId, headings, buf);
int level = m.group(1).length();
while (headings.size() >= level) headings.removeLast();
headings.addLast(m.group(2).strip());
continue;
}
if (buf.length() + para.length() > maxChars && buf.length() > 0) {
String tail = buf.substring(Math.max(0, buf.length() - overlapChars));
flush(out, docId, headings, buf);
buf.append(tail).append("\n\n");
}
buf.append(para.strip()).append("\n\n");
}
flush(out, docId, headings, buf);
return out;
}
private void flush(List<Chunk> out, String docId, Deque<String> h, StringBuilder buf) {
if (buf.toString().isBlank()) { buf.setLength(0); return; }
out.add(new Chunk(docId + ":" + out.size(), docId,
String.join(" > ", h), buf.toString().strip()));
buf.setLength(0);
}
}The heading path is embedded and indexed with the body, so a query about partial refunds matches a chunk whose body never repeats those words. Overlap carries the end of a long section into the next chunk so a sentence split by the size limit still appears whole in one of them. Measure chunk length in tokens against your embedding model's input limit; depending on the API, longer text is rejected or silently truncated.
Incremental ingestion
Re-embedding a whole corpus every night is slow and expensive, and it hides bugs because everything always looks fresh. Hash each chunk, compare with what is stored, and embed only what changed.
public void ingest(String tenantId, String docId, String version, String markdown) throws SQLException {
List<Chunk> chunks = chunker.split(docId, markdown);
Map<String, String> existing = loadHashes(tenantId, docId, embedder.modelId()); // chunk_id -> hash
List<Chunk> changed = new ArrayList<>();
for (Chunk ch : chunks) {
String h = sha256(ch.headingPath() + "\n" + ch.body());
if (!h.equals(existing.get(ch.id()))) changed.add(ch);
}
List<float[]> vectors = changed.isEmpty() ? List.of()
: embedder.embedDocuments(changed.stream()
.map(ch -> ch.headingPath() + "\n" + ch.body()).toList());
try (Connection cx = db.getConnection()) {
cx.setAutoCommit(false);
upsert(cx, tenantId, version, changed, vectors); // ON CONFLICT (tenant_id, chunk_id, embed_model)
deleteBeyond(cx, tenantId, docId, chunks.size()); // doc got shorter: drop stale tail chunks
cx.commit();
}
}Two details matter. Deleting chunks beyond the new count handles documents that got shorter; without it, the old tail stays searchable forever. And the whole document is replaced in one transaction, so a query never sees half of version 7 mixed with half of version 8. Deleted documents need an explicit tombstone from the source system, or a periodic reconciliation that lists source ids and removes rows whose document no longer exists. Ingestion that only ever adds is the most common cause of answers quoting retired policy.
Hybrid retrieval with reciprocal rank fusion
Vector search is good at meaning and bad at exact strings. A user asking about error code E-4021 or SKU 88-1173 needs keyword matching, because embeddings blur rare tokens. Run both searches and merge the rankings. Reciprocal rank fusion is the simplest robust merge: each result scores the sum of one over sixty plus its rank in each list, so items ranked well by both rise to the top and the raw scores, which are on incomparable scales, never need normalising.
WITH vec AS (
SELECT chunk_id, row_number() OVER (ORDER BY embedding <=> ?::vector) AS r
FROM kb_chunk
WHERE tenant_id = ? AND embed_model = ?
ORDER BY embedding <=> ?::vector
LIMIT 40
), kw AS (
SELECT chunk_id,
row_number() OVER (ORDER BY ts_rank_cd(tsv, q) DESC) AS r
FROM kb_chunk, websearch_to_tsquery('english', ?) AS q
WHERE tenant_id = ? AND embed_model = ? AND tsv @@ q
ORDER BY ts_rank_cd(tsv, q) DESC
LIMIT 40
), fused AS (
SELECT chunk_id,
COALESCE(1.0 / (60 + vec.r), 0) + COALESCE(1.0 / (60 + kw.r), 0) AS rrf
FROM vec FULL OUTER JOIN kw USING (chunk_id)
)
SELECT c.chunk_id, c.doc_id, c.heading_path, c.body, f.rrf
FROM fused f
JOIN kb_chunk c ON c.chunk_id = f.chunk_id AND c.tenant_id = ? AND c.embed_model = ?
ORDER BY f.rrf DESC
LIMIT ?;Pull more candidates than you return so fusion has something to work with. A cross-encoder reranker, if you add one, goes after fusion; keep it behind an interface so you can measure whether it earns its latency. The constant 60 comes from the original RRF paper.
The retrieval tool
Expose retrieval to the model as an instance function tool, so the retriever and its database connection are injected and can be mocked in tests. The tenant comes from session state that your application writes when it creates the session; it is never a tool parameter, because anything the model can set, a prompt injection can set too.
public final class DocSearchTools {
private final HybridRetriever retriever; // wraps the SQL above, injected, mockable
public DocSearchTools(HybridRetriever retriever) { this.retriever = retriever; }
@Schema(description = "Search the company knowledge base. Use for questions about "
+ "policies, products or procedures. Returns passages with ids; cite the ids you use.")
public Map<String, Object> searchDocs(
@Schema(name = "query", description = "A focused search query in plain words") String query,
ToolContext toolContext) {
Object tenant = toolContext.state().get("tenant_id"); // set by the app, never by the model
if (!(tenant instanceof String tenantId) || tenantId.isBlank()) {
return Map.of("status", "error", "error_code", "NO_TENANT",
"message", "Search is unavailable for this session.");
}
if (query == null || query.isBlank() || query.length() > 500) {
return Map.of("status", "error", "error_code", "INVALID_QUERY",
"message", "Send a short, specific search query.");
}
List<Hit> hits = retriever.search(tenantId, query, 6);
if (hits.isEmpty()) {
return Map.of("status", "success", "passages", List.of(),
"note", "No matching documents. Say so rather than guessing.");
}
return Map.of("status", "success", "passages", hits.stream().map(h -> Map.of(
"id", h.chunkId(), "section", h.headingPath(), "text", h.body())).toList());
}
}
LlmAgent agent = LlmAgent.builder()
.name("kb_assistant")
.model("<a model id your project can use>")
.instruction("Answer only from passages returned by searchDocs. Search before answering any "
+ "policy or product question; search again with different words if results look off-topic. "
+ "Cite passage ids in square brackets. If nothing relevant is found, say you do not know.")
.tools(FunctionTool.create(new DocSearchTools(retriever), "searchDocs"))
.build();Return only id, section and text per passage: the function response is stored in the session and re-sent on later turns. Compile with -parameters so ADK sees the parameter names and recognises toolContext by name, as the function tool guide explains.
Worked example
A support agent for tenant acme receives: "Can I refund only the shipping on order 5512?" The model calls searchDocs with the query "refund shipping only partial refund". The vector search ranks the chunk billing-faq:3 (heading path Billing > Refunds > Partial refunds) first and a returns-policy chunk second. The keyword search, matching "shipping" and "refund", ranks shipping-policy:2 first and billing-faq:3 third.
Fusion scores billing-faq:3 at 1/61 + 1/63, about 0.0323. shipping-policy:2 is not in the vector top 40, so it scores 1/61, about 0.0164. The tool returns six passages led by the partial-refund procedure. The model answers with the steps and cites [billing-faq:3], and the grounding callback checks that cited id was actually returned this turn. Had the tenant key been missing from state, the tool would have refused before running a single query.
Measure retrieval separately from answers
Answer quality mixes retrieval, prompt and model. Retrieval on its own is cheap to measure and is where most regressions start. Keep a golden set of fifty to two hundred real questions, each labelled with the chunk ids that answer it, and compute recall at k on every change to chunking, embeddings or the query.
// golden.jsonl: {"tenant":"acme","query":"can I refund part of an order","relevant":["billing-faq:3"]}
double recallAtK(List<GoldenCase> cases, HybridRetriever r, int k) {
int found = 0, total = 0;
for (GoldenCase g : cases) {
Set<String> got = r.search(g.tenant(), g.query(), k).stream()
.map(Hit::chunkId).collect(Collectors.toSet());
for (String rel : g.relevant()) { total++; if (got.contains(rel)) found++; }
}
return total == 0 ? 0 : (double) found / total;
}Chunk ids change on re-chunking, so label cases by document and heading and resolve ids at test time. Track recall at 6 for the final list and recall at 40 for each candidate list: if the candidates contain the answer but the final list does not, fix fusion or reranking; if neither does, fix chunking or the embedding model.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Answers quote a retired policy | Ingestion never deletes | Tombstones or periodic reconciliation |
| Exact codes and SKUs not found | Vector-only search | Hybrid search with RRF |
| Steps without context | Fixed-size chunks, no heading path | Structure-aware chunker with breadcrumbs |
| Another tenant's document appears | Tenant taken from a tool argument | Tenant only from session state, filtered in SQL |
| Fewer results than requested | Selective filter after approximate search | See the filtered-search trap; partition or raise candidates |
| Slow turns | Embedding call and two scans on the hot path | Bound with a timeout; cache query embeddings |
| Quality drops after model swap | Old and new vectors mixed | Filter by embed_model; backfill, then switch |
Put a timeout on the tool, as described in tool timeout handling, and return an error map the model can relay rather than letting the turn hang on a slow database.
What to do next
- Decide build or buy: try VertexAiRagRetrieval with Knowledge Engine first if your data and tenancy fit it.
- Create the kb_chunk table with both indexes and record the embedding model id on every row.
- Replace fixed-size chunking with a structure-aware chunker that carries heading paths.
- Make ingestion incremental with content hashes, and handle shorter and deleted documents explicitly.
- Add keyword search and reciprocal rank fusion, and expose retrieval as an instance FunctionTool that reads tenant from state.
- Build a golden set and gate every pipeline change on recall at k before looking at answer quality.