An ADK Java agent remembers a conversation through its session: the ordered list of events that the session service stores and replays into each model call. That memory ends when the session ends. Long-term memory, the kind that lets an agent recall next month that a user deploys to a particular region or prefers Gradle over Maven, lives behind a different interface, BaseMemoryService. The ADK Java API reference lists two implementations of it: InMemoryMemoryService, which the API reference describes as for prototyping only and which matches by keyword rather than meaning, and FirestoreMemoryService. If you want recall by meaning over a store you control, you implement the interface yourself on a vector store.
This article builds that implementation on PostgreSQL with the pgvector extension, because many Java teams already run Postgres and it keeps memory, sessions and access control in one operational surface. The concept of a memory service is introduced in the ADK Java memory service overview, and the choice between pgvector and dedicated vector databases is weighed in memory vector stores for agentic systems. Here the focus is the code and the failure modes: what goes in, how it is keyed, how it is searched, and what breaks.
The contract you are implementing
The interface is small. addSessionToMemory(Session session) returns an RxJava Completable and ingests a session; the API documentation notes that a session may be added multiple times during its lifetime. searchMemory(String appName, String userId, String query) returns Single<SearchMemoryResponse>, whose memories() is a list of MemoryEntry objects. Each entry carries a Content, an optional author and an optional timestamp, built with MemoryEntry.builder().
Two consequences follow before any vector appears. First, because the same session can arrive repeatedly, ingestion must be idempotent or every re-add duplicates memories and skews retrieval toward whatever was said in long sessions. Second, the search signature carries the application name and user ID, which tells you the scope ADK expects: memory belongs to one user of one app. Your store must enforce that scope on every read, because nothing else will.
On the read side, agents reach memory through a tool. LoadMemoryTool exposes a load-memory function to the model, which calls ToolContext.searchMemory(query) and gets back the current user's memories; its documentation notes that it currently uses only the text parts of each entry. So store text you want the model to read, and keep anything else as columns for yourself.
What to store: turns, chunks or extracted facts
There are three reasonable granularities, and the choice drives retrieval quality more than the choice of index.
| Unit | How it is made | Strength | Weakness |
|---|---|---|---|
| Raw events | Each final user or agent text event becomes one row | Cheap, faithful, no extra model calls | Noisy: greetings, tool chatter, restated questions |
| Chunks | Consecutive events merged to a size budget with overlap | Keeps question and answer together | Chunk boundaries split facts; more tokens per hit |
| Extracted facts | A model call turns a session into short, standalone statements | Dense, deduplicable, easy to correct or delete | Costs a model call per session; extraction can hallucinate |
A pragmatic path is to start with filtered raw events, which the code below does, measure retrieval on real questions, and add an extraction step when noise dominates. Whatever the unit, skip streaming fragments (events whose partial() is true), skip events that are only function calls or responses unless their results are meant to be remembered, and run a secret and PII filter before anything is embedded, because a vector store is a copy of user data that is easy to forget during deletion requests.
The schema
One table holds memory rows. The unique constraint is what makes repeated addSessionToMemory calls safe, and it includes the embedding model identifier so that a migration to a new model can write new rows alongside old ones.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE agent_memory (
id bigserial PRIMARY KEY,
app_name text NOT NULL,
user_id text NOT NULL,
session_id text NOT NULL,
event_id text NOT NULL,
author text,
occurred_at timestamptz NOT NULL,
body text NOT NULL,
embed_model text NOT NULL,
embedding vector(768) NOT NULL, -- must equal your model's dimension
UNIQUE (app_name, user_id, event_id, embed_model)
);
CREATE INDEX agent_memory_hnsw ON agent_memory
USING hnsw (embedding vector_cosine_ops);
CREATE INDEX agent_memory_scope ON agent_memory (app_name, user_id, occurred_at DESC);The cosine operator class matches the <=> distance operator used in queries. Pick one distance and use it in both the index and the query, or the index is ignored. HNSW in depth explains the graph the index builds and its recall-versus-latency knobs.
The implementation
The embedding call sits behind an interface of our own, not an ADK type, so the service can use Vertex AI, a Spring AI model or a local ONNX model without changing. Everything blocking runs on RxJava's I/O scheduler so JDBC and HTTP calls never block the agent's execution thread. Imports are omitted; the ADK types come from com.google.adk.memory, com.google.adk.sessions and com.google.adk.events, and Content and Part from com.google.genai.types.
public interface EmbeddingClient { // our seam, not an ADK API
String modelId(); // stored with every row
List<float[]> embedDocuments(List<String> texts);
float[] embedQuery(String text);
}
public final class PgVectorMemoryService implements BaseMemoryService {
private static final int MIN_CHARS = 20;
private final DataSource db;
private final EmbeddingClient embedder;
private final int topK;
private final double maxDistance; // cosine cut-off, tuned on your data
public PgVectorMemoryService(DataSource db, EmbeddingClient embedder, int topK, double maxDistance) {
this.db = db; this.embedder = embedder; this.topK = topK; this.maxDistance = maxDistance;
}
@Override
public Completable addSessionToMemory(Session session) {
return Completable.fromAction(() -> index(session)).subscribeOn(Schedulers.io());
}
private void index(Session s) throws SQLException {
List<Event> kept = new ArrayList<>();
List<String> texts = new ArrayList<>();
for (Event e : s.events()) {
if (e.partial().orElse(false)) continue; // streaming fragment
String t = textOf(e).strip();
if (t.length() < MIN_CHARS || Redactor.containsSecret(t)) continue;
kept.add(e); texts.add(t);
}
if (texts.isEmpty()) return;
List<float[]> vectors = embedder.embedDocuments(texts);
String sql = """
INSERT INTO agent_memory (app_name, user_id, session_id, event_id, author,
occurred_at, body, embed_model, embedding)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?::vector)
ON CONFLICT (app_name, user_id, event_id, embed_model) DO NOTHING
""";
try (Connection c = db.getConnection(); PreparedStatement ps = c.prepareStatement(sql)) {
for (int i = 0; i < kept.size(); i++) {
Event e = kept.get(i);
ps.setString(1, s.appName());
ps.setString(2, s.userId());
ps.setString(3, s.id());
ps.setString(4, e.id());
ps.setString(5, e.author());
ps.setTimestamp(6, Timestamp.from(Instant.ofEpochMilli(e.timestamp())));
ps.setString(7, texts.get(i));
ps.setString(8, embedder.modelId());
ps.setString(9, toLiteral(vectors.get(i)));
ps.addBatch();
}
ps.executeBatch();
}
}
@Override
public Single<SearchMemoryResponse> searchMemory(String appName, String userId, String query) {
return Single.fromCallable(() -> search(appName, userId, query)).subscribeOn(Schedulers.io());
}
private SearchMemoryResponse search(String app, String user, String query) throws SQLException {
String vec = toLiteral(embedder.embedQuery(query));
String sql = """
SELECT author, occurred_at, body, embedding <=> ?::vector AS distance
FROM agent_memory
WHERE app_name = ? AND user_id = ? AND embed_model = ?
ORDER BY embedding <=> ?::vector
LIMIT ?
""";
List<MemoryEntry> out = new ArrayList<>();
try (Connection c = db.getConnection()) {
c.setAutoCommit(false); // SET LOCAL needs a transaction
try {
try (Statement st = c.createStatement()) {
st.execute("SET LOCAL hnsw.iterative_scan = relaxed_order"); // pgvector 0.8.0+
}
try (PreparedStatement ps = c.prepareStatement(sql)) {
ps.setString(1, vec); ps.setString(2, app); ps.setString(3, user);
ps.setString(4, embedder.modelId()); ps.setString(5, vec); ps.setInt(6, topK);
try (ResultSet rs = ps.executeQuery()) {
while (rs.next()) {
if (rs.getDouble("distance") > maxDistance) continue;
out.add(MemoryEntry.builder()
.content(Content.fromParts(Part.fromText(rs.getString("body"))))
.author(rs.getString("author"))
.timestamp(rs.getTimestamp("occurred_at").toInstant())
.build());
}
}
}
c.commit();
} catch (SQLException | RuntimeException ex) {
c.rollback(); // never return a pooled connection mid-transaction
throw ex;
}
}
return SearchMemoryResponse.builder().memories(out).build();
}
private static String textOf(Event e) {
return e.content().flatMap(Content::parts).stream().flatMap(List::stream)
.map(p -> p.text().orElse("")).collect(Collectors.joining(" "));
}
private static String toLiteral(float[] v) { // pgvector text form: [0.1,0.2,...]
StringBuilder sb = new StringBuilder("[");
for (int i = 0; i < v.length; i++) { if (i > 0) sb.append(','); sb.append(v[i]); }
return sb.append(']').toString();
}
}Wiring takes two changes. Pass the service to the runner through Runner.builder() and its memoryService(...) method, and give the agent new LoadMemoryTool() in its tool list with an instruction saying when recall is worthwhile, for example before answering questions about the user's own environment or earlier decisions. Then decide when ingestion runs: at an explicit end-of-conversation hook, on idle timeout, or periodically for long-lived sessions. The idempotent upsert makes all three safe to combine. Redactor stands for whatever secret scanner you already use.
The filtered-search trap
Approximate indexes and WHERE clauses interact badly. HNSW search explores a candidate list hnsw.ef_search entries long, 40 by default, and the filter is applied to the rows it finds. If the current user owns 1 percent of the table, most of those 40 candidates belong to other users and are discarded, and the query can return one memory or none even though the user has dozens of relevant ones. The failure is silent: no error, just an agent that seems forgetful for some users and not others.
There are three remedies. pgvector 0.8.0 added iterative index scans: hnsw.iterative_scan set to relaxed_order or strict_order keeps scanning until enough rows pass the filter or hnsw.max_scan_tuples is reached, which is why the code sets it per transaction. Relaxed order can return rows slightly out of distance order; re-sort in Java if that matters. For a few very large tenants, partial indexes or partitions per tenant keep each index dense in rows that pass the filter. For small per-user memory sets, an exact scan over rows selected by the scope index is fast and has perfect recall. Check the plan with EXPLAIN ANALYZE for a small user and a large one.
Worked example: recall across sessions
On Monday a user of a deployment assistant says: "our payments service runs in eu-west-1 and we deploy with Argo CD." At session end, addSessionToMemory keeps that user event and the agent's confirmation, embeds both and inserts two rows scoped to the app and user. An idle-timeout job re-adds the session an hour later; the unique constraint turns the second insert into a no-op.
On Thursday, in a new session, the user asks "why is the payments rollout stuck?" The model calls the memory tool with a query such as "payments service deployment setup". The service embeds the query, filters to the user and ranks rows by cosine distance. In an illustrative run, not a benchmark, the Monday row scores a distance around 0.2, an unrelated row about a billing export around 0.5, and a cut-off of 0.35 keeps only the first. The model receives one short memory with its timestamp and asks about Argo CD sync status in eu-west-1 instead of making the user repeat their setup. The timestamp matters: when memories conflict, the instruction should tell the model to prefer the most recent and say which one it relied on.
Failure modes
| Failure | Cause | Prevention |
|---|---|---|
| Duplicate memories crowd results | Re-added sessions without a unique key | Unique (app, user, event, model) with ON CONFLICT DO NOTHING |
| One user sees another's memories | Missing scope filter, or a cache keyed only by query | Scope in every query; a test asserting zero cross-user rows |
| Recall collapses after a model upgrade | Old and new embeddings compared in one space | embed_model column, dual-write, backfill, then switch reads |
| Stale or contradictory facts returned | Nothing ever supersedes a memory | Timestamps in entries, recency-aware instructions, fact extraction with updates |
| Stored prompt injection | Untrusted text such as fetched pages or emails saved as memory | Remember only user and agent turns; record provenance; never store tool output blindly |
| Agent threads stall | JDBC or embedding HTTP on the execution thread | subscribeOn(Schedulers.io()) and timeouts on both calls |
| Deletion leaves data behind | Vectors not covered by retention jobs | Delete by user_id in the same job that clears sessions |
The poisoning row deserves emphasis. Memory turns a single successful injection into a persistent one: text saved today is replayed into prompts for months. Treat memory writes with the same suspicion as tool outputs, and see RAG for ADK Java agents for the parallel concerns on document retrieval.
Operating it
Measure retrieval, not just latency. Keep a small labelled set of (user, question, memory that should be found) triples, gathered with consent, and track recall at k and the fraction of searches returning nothing, split by how many memories the user has. Log the distance of every returned memory; a drifting distance distribution is the earliest sign of an embedding or data change. Budget latency as one embedding call plus one indexed query per recall, and cap how many recalls one turn may trigger.
For an embedding model change, write the new model's rows alongside the old, backfill from the stored body text, compare recall on the labelled set, then flip the model ID that reads use and delete the old rows. Because the body text is stored, re-embedding never needs the original sessions.
What to do next
- Prototype with
InMemoryMemoryServiceandLoadMemoryToolto confirm the agent recalls at the right moments. - Decide the unit: filtered raw events first, extracted facts once noise shows up in evaluation.
- Create the table with the unique constraint, the model column, an HNSW index and the scope index.
- Implement the service with I/O on
Schedulers.io(), a secret filter before embedding and a distance cut-off. - Enable iterative scans and check
EXPLAIN ANALYZEfor your smallest and largest users. - Write a test proving that one user's search never returns another user's rows.
- Build a labelled recall set and alert on recall and empty-result rate.
- Add memory rows to deletion and retention jobs, and plan the embedding-migration path before you need it.