An ADK Java agent spends most of its life waiting on a language model or a tool. It spends a few milliseconds per step writing an event to PostgreSQL. That ratio decides the size of the connection pool, and agents need surprisingly few connections. A pool sized for concurrent sessions rather than concurrent database work will exhaust the database long before it helps throughput.

This article sizes and operates the pool for adk-postgres, the site's own PostgreSQL implementation of ADK Java's BaseSessionService. (google/adk-java ships in-memory and Vertex AI session services, plus a Firestore one in contrib, but no JDBC service.) The schema and the appendEvent transaction, with its SELECT ... FOR UPDATE on the session row, are covered in the event log storage article.

Advertisement

Where adk-postgres borrows a connection

Agent invocations200 per pod, ~8 s eachModel + tool callsseconds, NO connectionmost of the timegetSession / appendEventBounded DB scheduler8 threads, queue 400borrowHikariCP poolmax 8, timeout 2 s~4 ms per appendPostgreSQL8 vCPU: ~16-20 useful activestatement_timeout, lock_timeoutx 6 pods (HPA max 20)pools multiplyPgBouncer (transaction)optional: 160 clients, 20 serversSize each pod's pool from DB hold time, then cap the fleet total at what the database can actually run in parallel
Connections are borrowed only for short database steps. Model and tool calls run with no connection held, a bounded scheduler caps how many threads can wait on the pool, and the fleet's total is capped at what PostgreSQL can execute in parallel.

createSession runs one INSERT. getSession reads the session row, merges user and app state and loads events, which is a few statements and the largest result set. appendEvent opens a transaction, locks the session row with FOR UPDATE, inserts the event with the next sequence number, merges the state delta and commits. listSessions and deleteSession are occasional. All of them run blocking JDBC wrapped in Single.fromCallable and subscribed on a scheduler, because the interface returns RxJava types.

A typical turn is one getSession and then an appendEvent for each event the runner emits: the user message, each model response, each function call and function response, and the final answer. Each append holds a connection for a few milliseconds; the turn takes seconds. A connection is in use for well under one percent of a turn.

Little's law: the arithmetic of a pool

Little's law says the average number of items in a system equals the arrival rate times the time each item spends there. For a pool, the items are borrowed connections: busy connections = DB calls per second x average hold time.

def pool_budget(concurrent_turns, turn_seconds, holds_ms, p99_factor, headroom=1.5):
    """Connections one pod needs, by Little's law: busy = arrival rate x hold time."""
    turns_per_s = concurrent_turns / turn_seconds
    hold_s = sum(holds_ms) / 1000.0                  # DB time per turn, all calls summed
    mean_busy = turns_per_s * hold_s
    return mean_busy, mean_busy * p99_factor * headroom

# one getSession (10 ms) + 8 appendEvent calls (4 ms each) per turn
mean, size = pool_budget(200, 8.0, [10] + [4] * 8, p99_factor=4)
print(f"mean busy {mean:.2f}, provision {size:.1f}")   # mean busy 1.05, provision 6.3 -> pool 8

Worked example: one pod runs 200 concurrent invocations with an 8-second average turn, so 25 turns per second arrive. Each turn holds connections for 10 + 8 x 4 = 42 ms. Mean busy connections: 25 x 0.042 = 1.05. Averages hide bursts, though. Hold times have a tail: a lock wait on a hot session, an autovacuum pass, a checkpoint, or a slow getSession on a long conversation. Measure p99 hold time and multiply. If the p99 is four times the mean and you add 50% headroom, you need about 6.3, so a pool of 8 per pod is generous. The default of 10 is not far off, but it is a coincidence. The point is that 200 concurrent agents do not need 200 connections, or even 50.

Advertisement

The database side: how many connections can PostgreSQL use?

Each PostgreSQL connection is a backend process. Throughput rises with active connections until the cores are busy, then falls as contention grows. HikariCP's pool-sizing guide gives a starting point for the total active connections a database can use: connections = (core_count * 2) + effective_spindle_count. Its example is a 4-core server with one disk, giving 9 and rounded to 10. Load-test it. A database with 8 vCPUs is happiest with somewhere around 16 to 20 concurrently active queries.

Now multiply by the fleet. Six pods with pools of 8 is 48 connections, but the HorizontalPodAutoscaler can take you to 20 pods, which is 160 connections. PostgreSQL's default max_connections is 100, a few slots are reserved for superusers, and migrations, admin sessions and replication need their own. Write the budget down: max pods x pool size + other clients is at most max_connections minus a margin. If the numbers do not fit, either shrink per-pod pools as the replica count grows, or put a pooler in front of the database.

HikariCP settings that matter

HikariConfig cfg = new HikariConfig();
cfg.setPoolName("adk-postgres");
cfg.setJdbcUrl("jdbc:postgresql://db.internal:5432/adk?ApplicationName=adk-agents");
cfg.setUsername("adk_service");
cfg.setPassword(System.getenv("ADK_DB_PASSWORD"));
cfg.setMaximumPoolSize(8);            // from the Little's-law budget below, not the default 10
cfg.setMinimumIdle(8);                // fixed-size pool; this is also HikariCP's default
cfg.setConnectionTimeout(2_000);      // fail an append in 2 s instead of the default 30 s
cfg.setMaxLifetime(25 * 60_000);      // several seconds below any proxy / LB / DB limit
cfg.setKeepaliveTime(60_000);         // ping idle connections so NATs do not drop them
cfg.setLeakDetectionThreshold(10_000);// log a stack trace if a borrow lasts > 10 s
cfg.setMetricRegistry(meterRegistry); // Micrometer: hikaricp.connections.* meters
DataSource ds = new HikariDataSource(cfg);
SettingDefaultWhat to choose for adk-postgres
maximumPoolSize10From the hold-time budget; usually 4-10 per pod
minimumIdlesame as maximumPoolSizeLeave it equal: a fixed pool avoids connect storms during spikes
connectionTimeout30000 ms (min 250 ms)1-3 s. A 30 s wait is a hung agent turn, not a slow one
maxLifetime1800000 msSeveral seconds shorter than any DB, proxy or load-balancer limit
keepaliveTime120000 ms (min 30000 ms)Below the shortest idle timeout of any NAT or firewall in the path
leakDetectionThreshold0 (off)On, above your slowest legitimate borrow, for example 10 s

The most consequential change is connectionTimeout. With the default, an exhausted pool makes every new append wait 30 seconds before failing, so a short database stall turns into a pile of frozen agent turns and client retries. A timeout of one to three seconds turns the same stall into fast, visible errors that the runner can surface or retry. Set ApplicationName in the JDBC URL so the service's connections are easy to find in pg_stat_activity.

Rule one: never hold a connection across a model or tool call

The biggest pool killer in agent code is not the pool size. It is a connection held while something slow happens. If a custom service, tool or callback keeps a transaction open across a model call, the connection is borrowed for seconds instead of milliseconds. Little's law then multiplies your requirement by a thousand. Worse, the session row stays locked by FOR UPDATE, so a concurrent append to the same session waits too.

// WRONG: one transaction spans the whole turn. The session row stays locked and the
// connection stays borrowed while the model thinks for 5-30 seconds.
try (Connection c = ds.getConnection()) {
    c.setAutoCommit(false);
    lockSession(c, sessionId);                 // SELECT ... FOR UPDATE
    LlmResponse r = model.generate(request);   // seconds of network I/O, connection idle-in-txn
    insertEvent(c, sessionId, toEvent(r));
    c.commit();
}

// RIGHT: talk to the model with no connection held, then append in one short transaction.
LlmResponse r = model.generate(request);
sessionService.appendEvent(session, toEvent(r)).blockingGet();   // borrow, lock, insert, commit: ms

The same rule applies to tools that query the database: borrow, query, return, then call the next API. If you need a guarantee across the model call, such as at-most-once tool execution, use an idempotency key or an outbox row rather than a held lock. The idempotency article covers that pattern. Set idle_in_transaction_session_timeout on the role, so a transaction that leaks open anyway is killed by the server instead of pinning a connection forever.

Rule two: bound the threads that wait on the pool

RxJava's Schedulers.io() is backed by a cached thread pool that grows without limit. That is fine while database calls are fast. When the pool is exhausted, every waiting appendEvent occupies an io() thread blocked inside getConnection(). A burst creates thousands of blocked platform threads. Give database work its own scheduler, sized to the pool, with a bounded queue:

// Schedulers.io() grows without bound: under load it creates one thread per waiting call,
// and each one blocks in HikariCP's getConnection() for up to connectionTimeout.
// Give database work its own bounded scheduler sized to the pool.
int poolSize = 8;
ThreadPoolExecutor dbExecutor = new ThreadPoolExecutor(
        poolSize, poolSize, 0L, TimeUnit.MILLISECONDS,
        new ArrayBlockingQueue<>(poolSize * 50),       // bounded backlog
        r -> { Thread t = new Thread(r, "adk-db"); t.setDaemon(true); return t; },
        new ThreadPoolExecutor.AbortPolicy());          // full queue = immediate, visible rejection
Scheduler dbScheduler = Schedulers.from(dbExecutor);

public Single<Event> appendEvent(Session session, Event event) {
    return Single.fromCallable(() -> { persist(session, event); return event; })
            .subscribeOn(dbScheduler)                   // not Schedulers.io()
            .flatMap(e -> BaseSessionService.super.appendEvent(session, e));
}

A full queue now rejects immediately, which is backpressure: the runner fails the turn with a clear error instead of hanging. If you run agents on virtual threads, the thread count stops being the problem, because a virtual thread waiting for a connection is cheap. Then the pool becomes the real concurrency limit, as the virtual threads article explains, and a Semaphore in front of the data source does the same bounding job.

One more threading trap is pool locking. If a call path borrows a second connection while holding the first, for example an outbox writer that calls ds.getConnection() again inside the append, enough concurrent threads can each hold one connection and wait forever for another. HikariCP's guide gives the minimum pool that avoids this deadlock as Tn x (Cm - 1) + 1, where Tn is the number of threads and Cm the connections each holds at once. The better fix is to pass the same Connection down, so Cm is 1.

Hot sessions and lock waits

FOR UPDATE serialises appends to one session, which is what keeps sequence numbers gap-free. It also means a session with concurrent writers turns lock waits into pool hold time: every waiting append holds a connection while it waits. Set lock_timeout (2 seconds, say) and statement_timeout on the service role, so a stuck session surfaces as an error within seconds. Watch lock-wait time as its own metric. A rise in lock waits with flat throughput points to a hot session or a retry storm, not an undersized pool, and adding connections makes it worse.

-- Timeouts on the service role, so every pooled connection inherits them.
ALTER ROLE adk_service SET statement_timeout = '5s';
ALTER ROLE adk_service SET lock_timeout = '2s';                       -- hot-session FOR UPDATE waits
ALTER ROLE adk_service SET idle_in_transaction_session_timeout = '15s'; -- kills a leaked open txn

-- Who is holding connections right now, and waiting on what?
SELECT state, wait_event_type, wait_event, count(*),
       max(now() - xact_start) AS oldest_txn
FROM pg_stat_activity
WHERE application_name = 'adk-agents'
GROUP BY 1, 2, 3
ORDER BY 4 DESC;

PgBouncer in transaction mode

When the fleet total does not fit max_connections, put PgBouncer between the pods and PostgreSQL in transaction pooling mode. Each client transaction is assigned a server connection only for its duration, so 160 client connections can share 20 server connections. The catch is that session state does not survive between transactions. The sibling articles already set the app name with SET LOCAL inside each transaction for row-level security, and that is safe. A plain SET, session advisory locks, LISTEN and temporary tables are not safe.

Prepared statements are the other trap. pgJDBC switches to server-side prepared statements after a statement has been executed a few times, controlled by prepareThreshold. Older PgBouncer versions cannot route those, so the classic fix is prepareThreshold=0 in the JDBC URL. PgBouncer 1.21 and later can track protocol-level prepared statements when max_prepared_statements is set. Check your PgBouncer version before choosing. Keep HikariCP anyway: it still bounds per-pod concurrency.

Metrics and how to read them

  • Pending threads (hikaricp.connections.pending). Persistently above zero means work is waiting; check hold times before resizing.
  • Acquire time (hikaricp.connections.acquire). Seconds at p99 mean exhaustion.
  • Usage time (hikaricp.connections.usage). This is the hold time from Little's law. A jump points to a slow query, a lock wait or a held transaction.
  • Timeouts (hikaricp.connections.timeout). Each is a failed agent step; alert on the rate.
  • Server side. Track pg_stat_activity counts by state, especially idle in transaction, plus lock waits and database CPU. If database CPU is high and pending is high, the database is the limit, and a bigger pool only adds contention.

The observability article covers how to attach database spans to the invocation trace tree, which is the fastest way to find the tool that holds a transaction across a model call.

Failure modes

SymptomLikely causeFix
Turns hang about 30 s, then failPool exhausted with the default connectionTimeoutDrop connectionTimeout to 1-3 s; find what holds connections
Thread count climbs into the thousandsSchedulers.io() threads blocked in getConnection()Bounded DB scheduler or semaphore
too many connections during scale-outPods x pool exceeds max_connectionsFleet budget, smaller pools, or PgBouncer
Many idle in transaction backendsTransaction held across a model or tool callShort transactions; idle_in_transaction_session_timeout
Errors about prepared statements behind PgBouncerServer-side prepares in transaction modeprepareThreshold=0 or max_prepared_statements

What to do next

  1. Count your database calls per turn and measure their hold times (mean and p99) with HikariCP's usage metric.
  2. Run the Little's-law budget for your peak concurrent turns and set maximumPoolSize from it, with minimumIdle equal to it.
  3. Write the fleet budget: maximum HPA replicas x pool size + other clients, against max_connections minus a margin. Add PgBouncer if it does not fit.
  4. Lower connectionTimeout to 1-3 s, set maxLifetime below every infrastructure limit, and turn on leak detection.
  5. Replace Schedulers.io() for JDBC work with a bounded scheduler or a semaphore.
  6. Set statement_timeout, lock_timeout and idle_in_transaction_session_timeout on the service role.
  7. Grep tools and callbacks for transactions that span a model or tool call, and refactor them into short ones.
Key takeaway: Size the adk-postgres pool from database work, not from sessions: busy connections equal DB calls per second times hold time, which for agents is usually a handful per pod. Keep HikariCP fixed-size with a short connectionTimeout and a maxLifetime below every infrastructure limit, never hold a connection or transaction across a model or tool call, bound the threads that wait on the pool, cap pods x pool size against max_connections (or use PgBouncer in transaction mode), and watch pending, acquire and usage metrics.