The adk-postgres store on this site keeps everything an ADK Java agent remembers: one row per session in adk_sessions, the ordered event log in adk_events, and user-scoped and app-scoped state in adk_user_state and adk_app_state. Lose it and every conversation starts from nothing. Restore it badly and something worse happens: agents resume conversations from a past that no longer matches the world their tools changed.

This article designs backup and restore for that store specifically. It uses standard PostgreSQL 17 features, WAL archiving, incremental base backups, point-in-time recovery and logical dumps, then adds what is particular to agent memory: checking the session-to-event invariant after a restore, restoring one tenant without rewinding the others, and dealing with the tool calls and deletions that happened after the recovery point. The SQL and Java are written for the schema in the event log article; they were not run against a live database here, so test them on your own.

Advertisement

What a backup must preserve

The schema has one invariant that matters more than any other. Each session row carries last_seq, and the append transaction inserts an event with the next sequence number and bumps last_seq in the same transaction. After any restore, last_seq must equal the highest seq in the session's events, and the retained sequence must have no gaps; retention removes whole sessions or partitions, never single events. If it does not, the next append either collides with an existing event or leaves a hole, and replaying the session reconstructs the wrong state.

Any backup taken as a single consistent snapshot of the whole database preserves this automatically: a physical backup with its WAL, or one pg_dump run, which reads everything in one transaction. What breaks it is anything that copies tables at different moments: per-table exports, separate dumps per tenant taken minutes apart, or storage-level snapshots of volumes that are not crash-consistent with each other. Also include the migration history table, so a restored database states its own schema version, and every partition if adk_events is partitioned.

Choose RPO and RTO before tools

Recovery point objective (RPO) is how much recent history you can lose. Recovery time objective (RTO) is how long agents can be down. For agent memory, lost history is user-visible: a customer repeats themselves, or an agent forgets it already issued a refund. A typical target is an RPO of about a minute and an RTO under an hour, which rules out nightly dumps as the primary mechanism and points to continuous WAL archiving.

Write the numbers down, because they drive cost: a one-minute RPO means archiving WAL within a minute even on a quiet database, and an RTO is only real once a drill has measured it.

Advertisement

The three mechanisms and what each is for

MechanismGives youDoes not give you
Streaming replicaFast failover after a host diesAny protection from mistakes: a bad DELETE replicates in milliseconds
Base backups plus WAL archiveRestore to any moment in the retention windowPortability across major versions or partial restores
Logical dump (pg_dump)A portable, inspectable copy; single-table or single-tenant extractionPoint-in-time recovery; restore of a large database is slow because indexes are rebuilt

Use all three for different jobs. The replica handles hardware failure. The archive handles everything that needs rewinding: a buggy release that corrupted state, an operator mistake, ransomware. The logical dump is the escape hatch for version upgrades and for pulling one tenant's data out of an otherwise healthy system.

Architecture

Backup architecture for the adk-postgres storeAgent podsADK Java runnersPrimaryadk_sessions, adk_eventsStandby replicaavailability, not backupJDBCstreaming WALWAL archiveevery segment, object storeBase backupsweekly full, daily incrementalLogical dumppg_dump -Fd, weeklyarchive_commandpg_basebackupRestore hostbase + incrementals combined, WAL replayed to target timerestore_commandThen: verify the session invariants, reconcile tool side effects, re-apply deletions, open traffic.
Physical backups and the WAL archive provide point-in-time recovery; the logical dump is the portable copy. The replica is not a backup.

Keep backups in a different account or project from the database, with write-once retention, so one compromised credential cannot delete both. Encrypt at rest, and remember that agent events often contain personal data, so the backup store is in scope for the same access controls as production.

Configuring the primary and the backup jobs

# postgresql.conf on the primary (PostgreSQL 17)
wal_level = replica
archive_mode = on
archive_command = '/opt/adk/bin/archive-wal %p %f'   # must exit 0 only after the copy is durable
archive_timeout = 60s                               # bounds RPO on a quiet database
summarize_wal = on                                  # required for incremental base backups

# weekly full and daily incremental base backups, run from a backup host
pg_basebackup -h db1 -U backup -D /backups/full-2026-09-27 -Fp -X none -c fast
pg_verifybackup -n /backups/full-2026-09-27      # -n: WAL is in the archive, not the backup
pg_basebackup -h db1 -U backup -D /backups/incr-2026-09-28 -Fp -X none -c fast \
    --incremental=/backups/full-2026-09-27/backup_manifest

# weekly logical dump: portable across major versions, parallel only in directory format
pg_dump -h db1 -U backup -d adk -Fd -j 4 -f /backups/logical-2026-09-27

Notes on each piece. archive_command runs for every completed WAL segment; it must return success only once the segment is durably stored, and a failing command makes WAL pile up on the primary until its disk fills, so alert on archiver failures. archive_timeout forces a segment switch on quiet systems so the RPO holds at night.

PostgreSQL 17 added incremental base backups. With summarize_wal on, a WAL summarizer process records which blocks changed, and pg_basebackup --incremental copies only those blocks relative to an earlier backup's manifest. Summarisation must have been running across the whole period since the parent backup, so turn it on before the first full backup of the chain. Restoring needs the full backup and every incremental after it, combined with pg_combinebackup. Many teams use a dedicated tool such as pgBackRest or Barman instead; the design in this article applies either way.

The logical dump uses directory format because pg_dump -j parallelism works only with -Fd, and pg_restore -j can then restore in parallel as well.

Restore runbook: point-in-time recovery

# 1. Rebuild a full data directory from the chain (oldest first)
pg_combinebackup -o /restore/data /backups/full-2026-09-27 \
    /backups/incr-2026-09-28 /backups/incr-2026-09-29 /backups/incr-2026-09-30
pg_verifybackup -n /restore/data                 # the combined directory has its own manifest

# 2. Tell recovery where WAL lives and where to stop
cat >> /restore/data/postgresql.conf <<'EOF'
restore_command = '/opt/adk/bin/fetch-wal %f %p'
recovery_target_time = '2026-10-01 14:02:00+00'
recovery_target_action = 'pause'
EOF
touch /restore/data/recovery.signal

# 3. Start on an isolated host with no agent traffic, then inspect before promoting
pg_ctl -D /restore/data start
psql -d adk -f verify_restore.sql
psql -c "SELECT pg_wal_replay_resume();"   # ends recovery at the target and promotes

Restore to a new, isolated host, never over the damaged primary, which is evidence you may need. Choose the target time just before the incident; if the cause was a single bad transaction you can instead use recovery_target_xid or recovery_target_lsn. Setting recovery_target_action = 'pause' stops replay at the target with the database readable, so you can run the checks below and, if the target was wrong, restart recovery from the base backup with a different one. Only after verification do you resume, promote, and point agents at the new host.

Time the whole sequence in drills: fetching the backup, combining, replaying WAL, verification, cut-over. That total, not the database's size, is your real RTO.

Verify before you open traffic

-- verify_restore.sql: each query must return zero rows

-- 1. last_seq must equal the highest retained event sequence for the session
--    (inner join: sessions whose events were all expired are checked by your retention job)
SELECT s.app_name, s.user_id, s.session_id, s.last_seq, e.max_seq
FROM adk_sessions s
JOIN (SELECT app_name, user_id, session_id, max(seq) AS max_seq
      FROM adk_events GROUP BY 1, 2, 3) e USING (app_name, user_id, session_id)
WHERE s.last_seq <> e.max_seq;

-- 2. no gaps inside the retained range of each session's sequence numbers
SELECT app_name, user_id, session_id
FROM adk_events
GROUP BY 1, 2, 3
HAVING max(seq) - min(seq) + 1 <> count(*);

-- 3. no events stamped after the recovery target (clock or target mistakes)
SELECT count(*) FROM adk_events
WHERE created_at > timestamptz '2026-10-01 14:02:00+00'
HAVING count(*) > 0;

After a clean point-in-time recovery all three queries should return nothing, because the restore is transaction-consistent. A non-empty result means something else is wrong: a restore assembled from mixed backups, a hand-edited table, or a bug in the append path that was already there before the incident, which is worth knowing either way. Also compare row counts and the newest created_at against what you expected from the target time.

Make the check impossible to skip by putting it in the application. A small gate, run by the agent service at startup when a restore flag is set, refuses traffic if sessions disagree with their logs:

/** Refuses to start agent traffic if the restored store breaks its invariants. */
public final class RestoreGate {
    private static final String DRIFT = """
        SELECT count(*) FROM adk_sessions s
        JOIN (SELECT app_name, user_id, session_id, max(seq) AS m
              FROM adk_events GROUP BY 1, 2, 3) e
          USING (app_name, user_id, session_id)
        WHERE s.last_seq <> e.m
        """;

    public static void check(DataSource ds, long maxDrift) throws SQLException {
        try (Connection c = ds.getConnection();
             Statement st = c.createStatement()) {
            c.setReadOnly(true);
            try (ResultSet rs = st.executeQuery(DRIFT)) {
                rs.next();
                long bad = rs.getLong(1);
                if (bad > maxDrift) {
                    throw new IllegalStateException(bad + " sessions disagree with their event log");
                }
            }
        }
    }
}

Restoring one tenant without rewinding the rest

Often only one user is damaged, for example by a buggy tool that wrote garbage into their state. Rewinding the whole database would discard everyone else's good history. Instead, restore the backup to a side instance at the target time, extract that user's rows from adk_sessions, adk_events and adk_user_state, and replace them on the primary inside a single BEGIN and COMMIT. Leave adk_app_state alone: it is keyed by app only and shared by every user, so touch it only when restoring a whole app. Do it with agent traffic for that tenant paused, so no append races the replacement.

Extract from the side instance in one transaction too, for example with a single pg_dump filtered later, or COPY queries inside one repeatable-read transaction, so the four tables agree with each other. Then run the invariant checks scoped to that tenant.

What a restore cannot undo

Rewinding the database does not rewind the world. Between the recovery target and the incident, agents called tools: they sent emails, created tickets, charged cards. After the restore, those calls are missing from the event log, so an agent that resumes the conversation may decide to make them again.

Two designs make this manageable. First, give every tool call an idempotency key derived from the session and the event that requested it, and have tools that change external systems honour it; a repeated call then becomes harmless. Second, record outgoing side effects in a separate outbox table or external system with its own retention, so after a restore you can list which ones happened after the target and either replay their results into the sessions or tell the affected users.

Deletions have the opposite problem. A user who asked to be forgotten after the recovery target is resurrected by the restore. Keep a deletion ledger outside the database's restore scope and re-apply it as a mandatory step after every restore, before traffic opens. The same reasoning bounds backup retention: data you are obliged to delete lives on in backups until they expire, so set the retention window with that obligation in mind.

Drills and monitoring

  • Restore the latest backup to a scratch host on a schedule, run the verification script, and record the measured RTO. An untested backup is a hope.
  • Alert on archiver failures, on the age of the newest archived WAL segment, and on the age of the newest base backup.
  • Run pg_verifybackup on every full backup after it is taken, and on the combined output of every drill, since that is the only way to exercise the incremental chain end to end.
  • Restore the logical dump into the next major PostgreSQL version occasionally, so an upgrade path is proven before you need it.
  • Keep the runbook next to the code, with real host names and commands, and rehearse it with whoever is on call.

Failure modes

  • Silent archive gap. One missing WAL segment stops point-in-time recovery at that point. Monitor continuity, not just success counts.
  • Broken incremental chain. Summarisation was off for part of the period, or an intermediate incremental was deleted, and the newer ones cannot be combined. Retention must remove whole chains.
  • Replica mistaken for backup. A truncated table is truncated everywhere within a second.
  • Inconsistent partial export. Tables copied at different times; the invariant query catches it.
  • Wrong target time zone. Always write recovery targets with an explicit offset.
  • Duplicate side effects after rollback. Agents redo tool calls lost from history. Idempotency keys and an outbox contain it.
  • Restored personal data. Deletions after the target come back. Re-apply the deletion ledger.

Related reading

The schema and append transaction come from Postgres event log storage for ADK Java; the session row is covered in the adk-postgres session table. Partitioned event tables and their maintenance are in partitioning for high volume, schema versions in migrations and schema versioning, and the WAL that all of this rests on in PostgreSQL WAL in depth.

What to do next

  1. Write down RPO and RTO targets for agent memory and get them agreed.
  2. Turn on WAL archiving with archive_timeout and summarize_wal, and alert on archiver failures.
  3. Schedule full and incremental base backups, verify fulls and combined chains with pg_verifybackup, and a weekly directory-format pg_dump.
  4. Add the invariant SQL and the startup gate to your restore procedure.
  5. Introduce idempotency keys for side-effecting tools and a deletion ledger outside the database.
  6. Run a full restore drill this month and record the measured recovery time.
Key takeaway: Back up the adk-postgres store with continuous WAL archiving plus full and PostgreSQL 17 incremental base backups for point-in-time recovery, a weekly logical dump for portability, and a replica only for failover. Restore to an isolated host, pause at the target, and verify that every session's last_seq matches its highest event sequence with no gaps before opening traffic. Restore single tenants from a side instance in one transaction. Plan for what a restore cannot undo: tool side effects after the target, which need idempotency keys and an outbox, and deletions, which need a ledger re-applied after every restore. Prove all of it with timed drills.