Once an ADK Java agent serves more than one customer, someone asks how much each customer may use. A rate limiter answers a different question: how fast. It smooths bursts so the provider's per-minute limits and your own capacity are not exceeded. A quota meters consumption over a long window, such as a day or a month, and usually reflects a commercial agreement: this tenant bought 50 million credits a month, this team may spend 20 million of them, this free-tier user gets 400 thousand a day.

This article builds quota management for ADK Java from the ground up. You will define units and scopes, build a ledger that reserves before each model call and settles afterwards, enforce it with an ADK plugin, handle provider-side quota errors separately, and work through a sizing example. Request-rate limiting is covered in the ADK Java rate limiter and in rate limiting a ParallelAgent fan-out; this page assumes those exist and deals only with period quotas.

Why quotas are not slow rate limits

It is tempting to implement quotas as a very slow token bucket. That breaks in three ways. First, quotas reset on calendar boundaries tied to billing, such as the first of the month in the tenant's time zone, not on a rolling refill. Second, quotas are hierarchical: a user's spend also counts against the team and the tenant, and any of the three can be the binding limit. Third, quota decisions are audited. When a customer disputes a block, you must show which counter was exhausted, by which calls, at what time. A bucket in memory cannot answer that.

Quotas also have a precision problem that rate limits do not. You do not know how many tokens an LLM call will use until it finishes: the output length is decided by the model, and an agent may loop through several model and tool calls for one user message. So the system cannot simply check a counter before a call. It must reserve an estimate, let the call run, and settle the difference when the actual usage arrives. That reserve-then-settle pattern is the core of everything below.

Units, scopes and windows

Decide three things before writing code.

  • Units. Raw tokens are the natural measure, but input, output, cached and reasoning tokens are priced differently by providers. Convert them to a single internal unit, called credits here, using weights from configuration. For example, credits = input + 4 x output + 0.25 x cached input. Take real weights from your provider's current price list and version them; never hard-code prices in the plugin.
  • Scopes. A typical hierarchy is tenant, then team, then user. Each scope has its own limit. A call is allowed only if every scope on its path has room. Store the scope path on the session at creation time so the plugin never has to look it up per call.
  • Windows. Each limit has a window: daily or monthly, aligned to a time zone. The counter key includes the window, such as quota:t-42:2026-10, so a new window starts with a fresh counter and old counters remain as the audit record.

Add one more choice per limit: hard or soft. A hard limit denies the call. A soft limit allows it, records an overage and notifies someone. Most paid tenants want soft limits with an overage allowance, because a blocked agent in the middle of a customer conversation is worse than a bill. Free tiers want hard limits.

Where enforcement lives in ADK Java

ADK Java gives you the right hook. A plugin extends BasePlugin and is registered on the runner with Runner.builder().plugins(...). Its callbacks run for every agent in the tree, which matters because a quota must count sub-agent and parallel-branch calls as well as the root agent's. Three callbacks do the work. beforeModelCallback(CallbackContext, LlmRequest.Builder) runs before each model call and may return a Maybe<LlmResponse>; returning a response skips the model call entirely. afterModelCallback(CallbackContext, LlmResponse) sees the response, including usageMetadata(). onModelErrorCallback runs when the call fails.

Quota enforcement around every model call: reserve before, settle after, release on errorClient requesttenant, team, userRunnerrunAsync, pluginsQuotaPluginbeforeModelCallback: reserveafterModelCallback: settle actual usageonModelErrorCallback: releaseLlmAgentmodel callModel APIprovider quotaallowedCanned replyquota exhausteddeniedQuota ledgeruser / team / tenant counters per windowreserve, settlePlan configlimits, weights, soft marginsUsage eventsbilling, reportsRate limits protect capacity per second; quotas meter consumption per day or month.They are separate mechanisms with separate stores and separate error messages.
The plugin sits on every model call in the agent tree. The ledger holds one counter per scope and window; plan configuration supplies limits and credit weights.

Keep per-agent callbacks for agent-specific behaviour. Quota is cross-cutting, and a plugin cannot be forgotten when someone adds a sub-agent.

The ledger: reserve, settle, release

The ledger is the only stateful part, and its contract is small. Reserve atomically checks every scope and increments all of them, or none. Settle replaces the reservation with the actual amount. Release cancels a reservation that produced no usage. Every reservation has an expiry so that a crashed process cannot hold credits forever.

public interface QuotaLedger {
  /** Atomically reserve {@code credits} on every scope in {@code path}, or deny. */
  Decision reserve(List<String> path, String window, long credits, Duration ttl);

  /** Replace a reservation with the actual usage (may be more or less than reserved). */
  void settle(Reservation r, long actualCredits);

  /** Cancel a reservation that produced no billable usage. */
  void release(Reservation r);
}

public record Reservation(String id, List<String> path, String window, long credits) {}

public sealed interface Decision {
  record Allowed(Reservation r) implements Decision {}
  record SoftOver(Reservation r, String scope) implements Decision {}
  record Denied(String scope, long used, long limit) implements Decision {}
}

For a single-region service, implement the ledger in a relational database with one row per scope and window, and do the reserve as one transaction that locks the rows in a fixed order (tenant, then team, then user) so that concurrent reservations cannot deadlock. For higher throughput, use Redis with a Lua script that checks and increments all keys in one atomic step, and stream every settle event to durable storage for the audit trail. In both cases, settle must be idempotent: key it by reservation id so a retried settle does not double-charge.

Settle can push a counter past its limit, because the actual usage exceeded the estimate. That is correct. The overshoot is bounded by one call's worth of usage per in-flight request, and the next reserve will be denied. Do not try to cancel a response after the fact; the tokens were already spent.

The quota plugin

Here is the plugin. It reads the scope path from session state, reserves a fixed estimate, returns a canned response on denial, and settles from usage metadata.

public final class QuotaPlugin extends BasePlugin {
  private final QuotaLedger ledger;
  private final CreditWeights weights;          // loaded from versioned plan config
  private final long reserveCredits;            // estimate per model call
  private final OverageAlerts overageAlerts = OverageAlerts.create();   // your notifier
  private final Map<String, Reservation> open = new ConcurrentHashMap<>();

  public QuotaPlugin(QuotaLedger ledger, CreditWeights weights, long reserveCredits) {
    super("quota");
    this.ledger = ledger;
    this.weights = weights;
    this.reserveCredits = reserveCredits;
  }

  private static String key(CallbackContext ctx) {
    // Parallel branches share an invocation id, so include branch and agent.
    return ctx.invocationId() + "/" + ctx.branch().orElse("") + "/" + ctx.agentName();
  }

  @Override
  public Maybe<LlmResponse> beforeModelCallback(CallbackContext ctx, LlmRequest.Builder req) {
    @SuppressWarnings("unchecked")
    List<String> path = (List<String>) ctx.state().get("quota_path");   // set at session creation
    String window = weights.windowFor(Instant.now());
    Decision d = ledger.reserve(path, window, reserveCredits, Duration.ofMinutes(10));
    return switch (d) {
      case Decision.Allowed a -> { open.put(key(ctx), a.r()); yield Maybe.empty(); }
      case Decision.SoftOver s -> { open.put(key(ctx), s.r()); overageAlerts.notify(s.scope()); yield Maybe.empty(); }
      case Decision.Denied x -> Maybe.just(LlmResponse.builder()
          .content(Content.fromParts(Part.fromText(
              "Your " + x.scope() + " usage limit for this period has been reached.")))
          .build());
    };
  }

  @Override
  public Maybe<LlmResponse> afterModelCallback(CallbackContext ctx, LlmResponse resp) {
    resp.usageMetadata().ifPresent(u -> {
      Reservation r = open.remove(key(ctx));
      if (r != null) {
        long credits = weights.credits(
            u.promptTokenCount().orElse(0),
            u.candidatesTokenCount().orElse(0),
            u.cachedContentTokenCount().orElse(0),
            u.thoughtsTokenCount().orElse(0));
        ledger.settle(r, credits);
      }
    });
    return Maybe.empty();
  }

  @Override
  public Maybe<LlmResponse> onModelErrorCallback(
      CallbackContext ctx, LlmRequest.Builder req, Throwable error) {
    Reservation r = open.remove(key(ctx));
    if (r != null) ledger.release(r);
    return Maybe.empty();                       // let the error propagate to retry logic
  }
}

Three details deserve attention. The canned response for a denial is returned from the plugin, so the model is never called and the user sees a clear message instead of an exception; keep the wording neutral and add a machine-readable flag in session state if your front end needs to show an upgrade prompt. The settle step only runs when usage metadata is present; with streaming, partial responses may arrive without it, so the reservation stays open until the response that carries usage arrives, and the ledger's expiry covers the case where it never does. And the token counts come back as Optional<Integer> values from GenerateContentResponseUsageMetadata, so treat a missing field as zero, not as an error.

Model callbacks only see model calls. If tools also cost money, such as a paid search API, meter them the same way in beforeToolCallback and afterToolCallback. As a backstop against runaway agent loops, also set the run's maximum model calls in RunConfig (the default is 500), which caps how many reservations one user message can trigger.

Provider quotas are a separate problem

Your quotas and the model provider's quotas are different things, and mixing them up produces confusing errors. Your quota is a business limit per customer. The provider's quota is a limit on your whole project, and when it is exhausted the Gemini API returns HTTP 429 with status RESOURCE_EXHAUSTED. In the Google GenAI Java SDK this surfaces as a com.google.genai.errors.ApiException whose code() is 429.

Handle the two separately. A provider 429 is transient for the customer: retry with backoff, shift to another region or model if you have one, and alert your operators. It must never be reported to the customer as "your quota is exhausted", and it must not consume their quota, which is why the plugin releases the reservation on error. Conversely, a denial from your own ledger is final for the window and should not be retried. Use distinct messages and distinct metrics, and you will save your support team a lot of confused tickets.

Size the provider quota by summing tenants' expected peak per-minute usage and keeping it below the provider's limit with headroom, using the measurements described in token counting across models.

Worked example: sizing a tenant plan

Take a tenant on a plan of 50 million credits per month, with one team capped at 20 million and each user capped at 400 thousand per day. Credits are input tokens plus four times output tokens. The support agent averages six model calls per user message: a router, three specialist calls and two summarisation calls. A typical call uses 3,000 input tokens and 400 output tokens, so it costs 3,000 + 1,600 = 4,600 credits. One message therefore costs about 27,600 credits, and a user's daily cap allows roughly 14 messages.

Now set the reservation estimate. Reserving the average, 4,600, means a long answer can overshoot. Reserving the worst case, input plus four times the configured maximum output tokens, is safe but locks up headroom: with 2,048 maximum output tokens, that is 3,000 + 8,192 = 11,192 credits per call. With six calls in flight across parallel branches, the user needs 67,000 credits of free headroom just to start a message. A practical middle ground is to reserve the 95th percentile measured from your own settle logs, around 7,000 credits here, and accept that the last message of a day may overshoot the user cap by a few thousand credits.

Finally, check the hierarchy. Sixty users at their full daily cap for 22 working days would spend 60 x 400,000 x 22 = 528 million credits, far above the team's 20 million, so the team limit binds first. Alert the tenant administrator at 80 percent of the team limit so the denial is never a surprise.

Failure modes

Quota systems fail in ways that look like billing bugs to customers.

  • Leaked reservations. A process dies between reserve and settle. Without expiry, the credits are locked until the window ends. Give every reservation a time to live and sweep expired ones.
  • Double counting on retries. A retry layer that re-runs a failed call reserves again. That is correct only if the first reservation was released; make release and settle idempotent by reservation id.
  • Unmetered paths. A team adds an agent built with a separate runner that does not include the plugin. Build runners through one factory that always adds it, and reconcile ledger totals against the provider's billing export monthly.
  • Ledger outage. If the ledger is unreachable, choose a policy in advance. Fail open with local accounting and later reconciliation suits paid tiers; fail closed suits free tiers. Never decide this in the incident.

Trade-offs

Precision costs latency: a ledger round trip per model call adds up for agents that make dozens of calls. A local pre-allocation scheme, where each replica leases a block of credits and reserves locally, removes most round trips at the cost of up to one block of overshoot per replica.

Hard limits protect margins but interrupt users mid-task; soft limits keep agents running but need billing for overages. Large reservations keep you under the limit but strand headroom; small ones are efficient but overshoot. Choose per plan, write the choice down, and expose it in the plan configuration rather than in code. Pair the plugin with the callback design patterns in ADK Java callbacks so that quota, logging and guardrails stay independent and testable.

What to do next

Use this checklist to add quota management to an ADK Java service.

  1. Write down units, credit weights, scopes, windows, time zones and hard or soft behaviour for each plan.
  2. Store the scope path in session state when each session is created.
  3. Implement the ledger with atomic multi-scope reserve, idempotent settle and release, and reservation expiry.
  4. Add the quota plugin to every runner through a single factory, and meter paid tools with tool callbacks.
  5. Set reservation sizes from the 95th percentile of measured per-call credits, and revisit them monthly.
  6. Separate provider 429 handling from customer quota denials in code, messages and metrics.
  7. Alert tenants at 80 percent of each limit, and reconcile ledger totals against provider billing every month.
Key takeaway: Treat quotas as metered consumption, not rate: convert tokens to versioned credits, enforce tenant, team and user limits per calendar window, and reserve an estimate before each model call in a plugin, settling from usage metadata afterwards. Keep provider 429s separate from customer denials, expire reservations, and reconcile against billing.