A shared agent service has one failure mode that no amount of model capacity fixes: a single user, script or runaway browser tab sends turns faster than anyone else, and everyone else's latency climbs. Per-user rate limiting means every authenticated user gets a bounded share of the service, whatever the others do. This article builds that limit for an ADK Java application. The limit lives in a plugin, the state lives in Redis, and there are three separate limits, because one number cannot express what you need to protect.

Three neighbouring problems have their own articles. Smoothing calls against the model provider's quota is covered in the ADK Java rate limiter. Monthly credit budgets with reserve-and-settle accounting are in quota management, and capping the fan-out of a ParallelAgent is in rate limiting parallel agents. This page is about the end user: how fast one person may start turns, how many tokens per minute their turns may burn, and how many of their runs may be in flight at once.

Three limits, three questions

A turn in ADK Java is one call to Runner.runAsync with a user message. That turn can trigger one model call or fifty, depending on how many tools the agent calls and how many sub-agents it delegates to. A limit that counts only HTTP requests therefore misses most of the cost, and a limit that counts only tokens lets a user pile up dozens of concurrent runs that each look cheap at first. So you enforce three limits, each answering a different question:

LimitQuestion it answersUnitWhere it is checked
Turn rateHow often may this user start work?turns per minute, with a burstbeforeRunCallback, once per turn
Token rateHow much model capacity may this user consume?prompt plus output tokens per minutebeforeModelCallback and afterModelCallback, every model call
ConcurrencyHow many runs may this user have open at once?in-flight invocationsacquired in beforeRunCallback, released in afterRunCallback

The turn rate stops flooding, the token rate stops the user who sends few but enormous requests, and the concurrency cap stops the user with ten open tabs, each run holding memory and connections. As a backstop, RunConfig.maxLlmCalls caps model calls per run; its default in the current source is 500.

Where the check belongs in ADK Java

ADK Java has several callback points. Read the plugin interface in the adk-java repository rather than guessing, because two hooks that look interchangeable behave differently. The relevant signatures in com.google.adk.plugins.Plugin are:

Maybe<Content>     onUserMessageCallback(InvocationContext ctx, Content userMessage);
Maybe<Content>     beforeRunCallback(InvocationContext ctx);
Completable        afterRunCallback(InvocationContext ctx);
Completable        onRunErrorCallback(InvocationContext ctx, Throwable error);
Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder request);
Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse response);

If onUserMessageCallback returns content, that content replaces the user message and the run continues. It cannot stop the run, so it is the wrong place to deny. beforeRunCallback is the right place. In Runner.java (main branch, October 2026), content it returns is wrapped in an event authored by model and emitted instead of running the agent. afterRunCallback still runs afterwards, so your release logic must tolerate a run that never acquired anything. Note that the user's message event is created and appended to the session before beforeRunCallback fires, so a denied turn still leaves the user's message in the history; the model will see it on the next allowed turn.

Identity comes from InvocationContext.userId() (and userId() on the callback context). That value is whatever your server passed to runAsync. ADK does not authenticate it, so it must come from a verified token or session cookie on your side, never from a field in the request body. A limiter keyed on a client-supplied user id does nothing against an abusive user, who will simply send a fresh id with each request.

Plugins registered on the runner apply to every agent in the tree, which is what you want: a sub-agent's model calls count against the same user (per-agent hooks are covered in ADK Java callbacks). Register them through Runner.builder().plugins(...) or through an App's plugin list, but not both. The builder throws IllegalStateException if you combine app() with plugins().

Per-user limits inside one ADK Java runClientbearer tokenAuth layerverified userIdRunner.runAsyncuserId, sessionIdonUserMessageCallbackmay rewrite, cannot haltbeforeRunCallbackturn bucket + leasebeforeModelCallbacktokens-per-minute bucketafterModelCallbackreconcile actual tokensafterRunCallbackrelease leaseallowedmodel callturn endsCanned model eventagent skippeddeniedWait or denybounded delayover budgetRedisbuckets, leases
Turn admission and the concurrency lease are checked once per turn; the token bucket is checked around every model call, including sub-agent calls. All state is in Redis so every replica sees the same buckets.

A token bucket in Redis

Each per-user limit is a token bucket. A bucket holds up to capacity tokens and refills continuously at rate tokens per second. A request costing n first tops the bucket up with rate * elapsed (capped at capacity), then subtracts n. Capacity sets the burst a user may spend at once; rate sets the long-run average. For turns the cost is 1. For tokens the cost is an estimate made before the model call and corrected afterwards.

The token bucket is allowed to go negative, up to a limit. If a model call needs 6,000 tokens and the bucket holds 2,000, the script still subtracts the full 6,000 and returns how long the caller must wait until the balance is back to zero: here, 4,000 tokens at the refill rate. If that wait would exceed max_wait_ms, the script takes nothing and returns -1. Because the caller waits before acting, the long-run rate still holds, and a call larger than the whole bucket can still run.

The script runs atomically in Redis, so replicas never race, and it reads the clock with TIME, so application hosts cannot disagree about time. Redis 7 replicates script effects, which makes TIME inside a script safe.

-- KEYS[1] = bucket key, e.g. "rl:tok:{user-42}"
-- ARGV: rate (tokens per ms), capacity, cost, max_wait_ms
local t = redis.call('TIME')
local now = tonumber(t[1]) * 1000 + math.floor(tonumber(t[2]) / 1000)
local rate, cap = tonumber(ARGV[1]), tonumber(ARGV[2])
local cost, max_wait = tonumber(ARGV[3]), tonumber(ARGV[4])
local s = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(s[1]) or cap
local ts = tonumber(s[2]) or now
tokens = math.min(cap, tokens + (now - ts) * rate)
local after = tokens - cost
local wait = 0
if after < 0 then wait = math.ceil(-after / rate) end
if wait > max_wait then                      -- refuse: take nothing
  redis.call('HSET', KEYS[1], 'tokens', tostring(tokens), 'ts', now)
  redis.call('PEXPIRE', KEYS[1], math.ceil((cap - tokens) / rate) + 60000)
  return -1
end
redis.call('HSET', KEYS[1], 'tokens', tostring(after), 'ts', now)
redis.call('PEXPIRE', KEYS[1], math.ceil((cap - after) / rate) + 60000)
return wait

The braces in the key are a Redis Cluster hash tag. They put a user's turn bucket, token bucket and lease set in the same slot, so one script can touch all three. The concurrency lease is a sorted set per user whose members are invocation ids, scored by expiry time. Acquiring a lease runs ZREMRANGEBYSCORE to drop expired leases, then ZCARD to count, then ZADD if the count is under the cap. Releasing a lease is a single ZREM. The expiry matters: if a client disconnects and the reactive stream is disposed, afterRunCallback never runs, and without the expiry that slot would leak forever.

The plugin

The plugin ties the hooks to the store. RateStore wraps the two scripts and PlanLimits maps a user to their plan's numbers from configuration you can reload. One plugin instance serves every invocation concurrently, so its only local state is a concurrent map of lease holders.

public final class PerUserLimitPlugin extends BasePlugin {
  private final RateStore store;          // Redis scripts above
  private final PlanLimits plans;         // userId -> Limits, cached
  private final Map<String, String> leases = new ConcurrentHashMap<>();  // invocationId -> userId

  public PerUserLimitPlugin(RateStore store, PlanLimits plans) {
    super("per_user_limits");
    this.store = store;
    this.plans = plans;
  }

  @Override
  public Maybe<Content> beforeRunCallback(InvocationContext ctx) {
    String user = ctx.userId();
    Admission a = store.admit(user, ctx.invocationId(), plans.forUser(user));  // bucket + lease
    if (!a.allowed()) {
      return Maybe.just(Content.fromParts(Part.fromText(a.message())));
    }
    leases.put(ctx.invocationId(), user);
    return Maybe.empty();
  }

  @Override
  public Completable afterRunCallback(InvocationContext ctx) { return release(ctx); }

  @Override
  public Completable onRunErrorCallback(InvocationContext ctx, Throwable error) { return release(ctx); }

  private Completable release(InvocationContext ctx) {
    String user = leases.remove(ctx.invocationId());      // null when the turn was denied
    return user == null ? Completable.complete()
        : Completable.fromAction(() -> store.releaseLease(user, ctx.invocationId()));
  }

  @Override
  public Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder req) {
    Limits l = plans.forUser(ctx.userId());
    long waitMs = store.reserveTokens(ctx.userId(), l.estimatePerCall(), l);
    if (waitMs == 0) return Maybe.empty();
    if (waitMs > 0) return Completable.timer(waitMs, TimeUnit.MILLISECONDS).andThen(Maybe.<LlmResponse>empty());
    return Maybe.just(LlmResponse.builder()
        .content(Content.fromParts(Part.fromText(
            "You have reached your per-minute usage limit. Please try again in a minute.")))
        .build());
  }

  @Override
  public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
    if (resp.partial().orElse(false)) return Maybe.empty();     // settle once, on the final response
    resp.usageMetadata().ifPresent(u -> {
      long actual = u.promptTokenCount().orElse(0) + u.candidatesTokenCount().orElse(0);
      Limits l = plans.forUser(ctx.userId());
      store.adjustTokens(ctx.userId(), actual - l.estimatePerCall(), l);   // may be negative
    });
    return Maybe.empty();
  }

  @Override
  public Maybe<LlmResponse> onModelErrorCallback(
      CallbackContext ctx, LlmRequest.Builder req, Throwable error) {
    Limits l = plans.forUser(ctx.userId());
    store.adjustTokens(ctx.userId(), -l.estimatePerCall(), l);   // refund: no tokens were spent
    return Maybe.empty();                                        // let retry logic see the error
  }
}

Three choices are deliberate. First, the turn check denies immediately while the token check waits. A user who has just typed a message prefers a clear refusal to a frozen screen. A run that is already halfway through a tool chain, though, is better delayed by a second or two than abandoned with half its work done. Second, Completable.timer delays on an RxJava scheduler instead of blocking a thread, so a thousand waiting runs cost timers rather than threads. Third, the correction in afterModelCallback can be negative, which refunds an overestimate, or positive, which charges for an underestimate. Partial streamed chunks are skipped, so each call settles once; a failed call is refunded in onModelErrorCallback so provider errors do not drain the user's bucket.

A mid-turn denial ends the turn with a model reply, and the agent stops there. If your agent was in the middle of a multi-step tool plan, the user will see a partial answer followed by the limit message. Write that message to say plainly that the work stopped early.

Worked example: sizing a plan

Take a support agent with this plan: 20 turns per minute with a burst of 5, 60,000 tokens per minute with a burst equal to one minute's allowance, 2 concurrent runs, a per-call estimate of 6,000 tokens and a maximum mid-turn wait of 2 seconds. The turn bucket refills at 20/60 = 0.33 tokens per second, which is one turn every 3 seconds.

  1. A script fires 10 turns within one second. The first 5 drain the burst. By the 6th, about 0.33 tokens have refilled, which is not enough for 1, so turns 6 to 10 are denied. The retry hint is (1 - 0.33) / 0.33, about 2 seconds.
  2. A human user sends one turn every 10 seconds. That is 6 per minute, well under 20. The turn bucket never empties, and this user never notices the limit.
  3. A power user's turns average 3 model calls of 6,000 tokens: 18,000 tokens per turn. The full 60,000-token bucket covers 3 such turns at once, and the token limit binds long before the turn limit. On an empty bucket a 6,000-token call must wait 6 seconds at 1,000 tokens per second, over the 2-second maximum, so the call is refused mid-turn: this user sees turns cut short, not slowed. To make them wait instead, raise the maximum wait to about 6 seconds, or admit a turn only when the bucket holds a full turn's estimate.
  4. The same user opens a third tab while two long runs are still active. The lease set holds 2 live invocation ids, so the third turn is denied with a concurrency message, even though both buckets have room.

Derive the numbers from data: take the 95th percentile of turns and tokens per minute from a week of usageMetadata, set limits two to three times higher, and run the plugin in shadow mode (decide and log, always allow) before enforcing.

Operational guidance

  • Fail open or closed, on purpose. If Redis is unreachable, decide in advance. Interactive products usually fail open with a local in-memory bucket as a degraded fallback. Expensive batch agents usually fail closed. Either way, emit a metric and alert on the fallback path.
  • Tell the client early. Repeat the turn check at the HTTP layer so it can return 429 with Retry-After; the plugin stays the authority for model-call budgets.
  • Metrics per decision. Count allowed, waited and denied decisions by limit type and plan; record wait time as a histogram; export the top users by tokens per minute. Never put user ids in metric labels; keep them in logs.
  • Config as data. Keep plan limits in versioned configuration with a per-user override for incidents, and reload it without a deploy.

Failure modes

  • Client-chosen user id. If the limiter trusts an id the client supplies, it enforces nothing. Derive userId from the verified principal.
  • Denying in onUserMessageCallback. Returning content there replaces the message and the agent answers your denial text as if the user had typed it.
  • Leaked leases. A disposed stream skips afterRunCallback. Without lease expiry, a user who reloads the page twice is locked out for good.
  • Double release. afterRunCallback runs even for denied turns. Releasing by user instead of by invocation id frees someone else's slot.
  • Estimate drift. A fixed per-call estimate that is far too low lets bursts through and then charges heavy debt afterwards. Estimate from the request size, or track a moving average per agent.
  • Fan-out multiplication. A ParallelAgent with eight branches makes eight model calls at once, and all eight reserve tokens against one user's bucket. Size the bucket's burst for your widest fan-out, or the branches will throttle each other.

What to do next

  1. Check where your userId comes from today and make it the verified principal.
  2. Export a week of per-turn token usage and derive the three limits from the 95th percentile.
  3. Load the Lua script into Redis and unit-test it with a fake clock: burst, refill, the debt path and the -1 refusal.
  4. Register the plugin in shadow mode and dashboard the would-deny counts per plan.
  5. Kill a client mid-run in a test and confirm its lease expires and the slot comes back.
  6. Turn on enforcement for one plan, add the matching 429 check at the HTTP edge, then roll it out to the other plans.
Key takeaway: Per-user limiting in ADK Java is three limits: turns per minute, tokens per minute and concurrent runs. Deny turns in beforeRunCallback, whose content skips the agent; onUserMessageCallback only rewrites the message. Meter tokens around every model call with an atomic Redis bucket that can wait briefly. Hold concurrency as expiring leases, and key all of it on an identity you verified, not one the client sent.