Every multi-tenant platform eventually has to say no. A customer may not create a 501st virtual machine, a free-tier key may not make its millionth call this month, a team may not hold more than 64 GPUs. The component that says no is the quota system, and it is not the same thing as a rate limiter. A rate limiter protects a server from bursts over seconds; a quota system enforces a business or capacity contract over hours, months or the lifetime of a resource, across every server and region, and has to be auditable when a customer disputes it.

This article designs one from first principles: the three kinds of quota, the data model for a tenant hierarchy, how to enforce counts of things transactionally, how to enforce high-volume budgets without a network call per request, how to reconcile against metered usage, and what fails in production.

Advertisement

Three kinds of quota, three different mechanisms

Most quota bugs come from treating all limits alike. They split into three kinds, and each needs a different consistency model:

KindExampleWhat is countedCorrect mechanism
Rate100 writes per second per projectEvents in a short windowToken bucket at the edge; small overshoot is harmless
Budget (consumption)1,000,000 API calls per day; $5,000 of inference per monthEvents in a long window, often weighted by costLeases handed to frontends, reconciled against metering
AllocationAt most 500 VMs or 64 GPUs per projectThings that currently existTransactional reserve/commit/release against a strongly consistent store

The distinction matters because the cost of error differs. Overshooting a per-second rate by 3 percent costs nothing. Overshooting an allocation quota by one GPU may mean a capacity plan is violated or a physical host is oversubscribed. Budget quotas sit in between: a small, bounded overshoot is normally acceptable if it is reconciled and billed correctly. Google Cloud draws the same line between rate quotas and allocation quotas, and Kubernetes ResourceQuota is an allocation quota enforced at admission time.

The rate case is well covered by the existing rate limiting algorithms article, so the rest of this page focuses on allocation and budget quotas.

Requirements and the tenant hierarchy

A realistic set of requirements: limits at several levels (organisation, project, user or API key); per-resource limits that can be raised per customer without a deploy; enforcement across all regions; admission decisions in well under a millisecond on the hot path for budget quotas; no double counting when clients retry; an audit trail explaining every rejection; and a way for customers to see their usage and request more.

Model tenants as a tree. A request is admitted only if every node on the path from its leaf to the root has room. Child limits are allowed to sum to more than the parent, which is deliberate oversubscription: ten projects may each be allowed 100 GPUs while the organisation is allowed 300. The data model below keeps limits, usage and reservations in separate tables so that a limit change never rewrites usage and usage is always explainable from reservations.

-- One row per (scope, resource). Scope is a node in the tenant tree.
CREATE TABLE quota_limit (
  scope_id     TEXT NOT NULL,      -- 'org:acme', 'project:acme-prod', 'user:42'
  parent_id    TEXT,               -- NULL for the root
  resource     TEXT NOT NULL,      -- 'vm.count', 'gpu.count', 'api.calls.day'
  kind         TEXT NOT NULL,      -- 'allocation' | 'rate' | 'budget'
  limit_value  BIGINT NOT NULL,
  window_sec   INT,                -- NULL for allocation quotas
  PRIMARY KEY (scope_id, resource)
);

CREATE TABLE quota_usage (
  scope_id   TEXT NOT NULL,
  resource   TEXT NOT NULL,
  window_id  BIGINT NOT NULL,      -- 0 for allocation, floor(now/window) otherwise
  committed  BIGINT NOT NULL DEFAULT 0,
  reserved   BIGINT NOT NULL DEFAULT 0,
  PRIMARY KEY (scope_id, resource, window_id)
);

CREATE TABLE quota_reservation (
  reservation_id TEXT PRIMARY KEY, -- caller-supplied idempotency key
  scope_path     TEXT[] NOT NULL,  -- leaf to root
  resource       TEXT NOT NULL,
  amount         BIGINT NOT NULL,
  state          TEXT NOT NULL,    -- 'held' | 'committed' | 'released'
  expires_at     TIMESTAMPTZ NOT NULL
);
Advertisement

The architecture

The design has a synchronous path and an asynchronous one. Synchronously, service frontends call a regional quota service, either per operation (allocation) or per lease (budget). Asynchronously, every consumed unit is emitted as a usage event to a metering pipeline that builds the authoritative ledger used for billing and for correcting drift in the quota store. Limits and overrides come from an admin plane with an approval workflow.

ClientsAPI calls, create VMService frontendsquota client + lease cacheQuota service (regional)limits, leases, reservationsQuota storestrongly consistent rowsUsage eventslog / streamMetering + reconciliationledger, billing, drift reportAdmin + increase flowoverrides, approvalsrequestlease / reserveemit usagecorrectset limitsThe hot path touches only the frontend's lease cache; the quota service is consulted per lease, not per request.Allocation quotas (counts of things) always go to the store transactionally.
Quota system architecture: frontends admit from local leases, the quota service owns limits and reservations, and metering reconciles what was actually consumed.

Allocation quotas: reserve, commit, release

Creating a VM is not instant. The request is admitted, capacity is found, disks are built, and any step can fail. If you increment usage at admission and the create fails, usage leaks. If you increment only after success, two concurrent creates can both pass the check. The fix is a two-phase protocol: reserve before doing the work, commit when the resource exists, release if it does not. Reservations carry a time-to-live so a crashed caller cannot hold quota forever, and a caller-supplied reservation ID makes retries idempotent, as in any idempotent API.

def reserve(txn, reservation_id, scope_path, resource, amount, ttl_s=300):
    # Idempotent: a retried call with the same key returns the first answer.
    existing = txn.get_reservation(reservation_id)
    if existing:
        return existing.state != "released"

    # Lock every level in a FIXED order, root first, to avoid deadlock.
    rows = [txn.lock_usage(scope, resource, window_id=0)
            for scope in reversed(scope_path)]
    for row in rows:
        limit = txn.limit(row.scope_id, resource)
        if row.committed + row.reserved + amount > limit:
            return False              # caller maps this to RESOURCE_EXHAUSTED
    for row in rows:
        row.reserved += amount
    txn.insert_reservation(reservation_id, scope_path, resource, amount,
                           state="held", expires_at=now() + ttl_s)
    return True

def commit(txn, reservation_id):
    r = txn.lock_reservation(reservation_id)
    if r.state != "held":
        return r.state == "committed"  # idempotent
    for scope in reversed(r.scope_path):
        row = txn.lock_usage(scope, r.resource, 0)
        row.reserved -= r.amount
        row.committed += r.amount
    r.state = "committed"
    return True

# release() mirrors commit() but only decrements `reserved`.
# A sweeper calls release() on every 'held' row past expires_at.

Deletion decrements committed through the same kind of idempotent operation, keyed by the resource ID so a replayed delete cannot decrement twice. Because allocation traffic is low (people create VMs, not millions per second) a single strongly consistent store per region, or a globally consistent one, handles it comfortably. Hot rows appear only at the root of a very large organisation; if that happens, shard the root counter into N sub-counters with a limit of L/N each and rebalance headroom between them periodically.

Budget quotas at high volume: leases

A daily API budget cannot be checked with a database transaction per call: at 50,000 requests per second that is 50,000 contended writes per second on one tenant's row. Instead, the quota service hands each frontend a lease: permission to admit N more units for the next T seconds. The frontend admits locally from its lease and asks for more when it runs low.

class LeaseCache:
    # Per-frontend admission for high-volume rate/budget quotas.
    def __init__(self, client, tenant, resource):
        self.client, self.tenant, self.resource = client, tenant, resource
        self.remaining, self.expires = 0, 0

    def admit(self, cost=1):
        if self.remaining < cost or now() >= self.expires:
            self._refill()
        if self.remaining >= cost:
            self.remaining -= cost
            return True
        return False

    def _refill(self):
        # Ask for roughly 10 seconds of this frontend's recent demand.
        want = max(1, int(self.observed_rate() * 10))
        grant = self.client.lease(self.tenant, self.resource, want)  # may be < want
        self.remaining, self.expires = grant.amount, grant.expires_at
        # Unused tokens on expiry are returned (best effort) by a background task.

The quota service tracks outstanding leases as reserved usage. The maximum overshoot is therefore bounded by what has been leased but not yet reported, and you can tune it directly: smaller leases mean more calls to the quota service and less overshoot. When a tenant is near its limit, the service should shrink grants (for example, never grant more than a fraction of the remaining headroom divided by the number of active frontends) so that the last few percent of the budget is handed out in small pieces. Frontends report actual consumption when a lease expires, and unused units return to the pool.

Worked example: sizing leases for a daily budget

A tenant has a budget of 10,000,000 calls per day and is served by 40 frontends. It averages 80 calls per second overall, peaking at 400. Per frontend that is 2 calls per second on average and 10 at peak.

With a 10-second lease sized to recent demand, each frontend holds at most about 100 units at peak, so the fleet has at most 4,000 units outstanding. That is the worst-case overshoot if every lease is spent in the final moments of the day: 0.04 percent of the budget. The quota service sees 40 frontends refreshing every 10 seconds, or 4 calls per second for this tenant, instead of 400.

Near the end of the budget the service switches to headroom-proportional grants. With 2,000 units left it grants at most 2,000 / 40 / 2 = 25 per request, so the remaining budget is spread instead of captured by whichever frontend asked first. If the business requires zero overshoot, budget quotas must fall back to per-request transactions for the final slice, which is a latency and availability cost the product owner should choose knowingly.

Windows, metering and reconciliation

Long-window quotas need an explicit calendar. Decide whether a day is UTC or the customer's billing timezone and whether a month is calendar or rolling, and store the window identifier with the usage row so a reset is simply a new row rather than a destructive update at midnight. Resets done by a cron job fail in exactly the way you would expect when the job is late.

The quota store is an admission-control estimate; the metering ledger is the truth. Every served unit emits a usage event with a unique ID, and a pipeline deduplicates and aggregates those events, much as in a metrics aggregation system. A reconciliation job compares the ledger with the quota store per tenant and window, corrects the store, and reports drift. Persistent drift in one direction usually means a code path that consumes without emitting, or that emits twice on retry.

Failure modes

FailureWhat happensMitigation
Quota service unreachableFrontends cannot refresh leasesPer-quota policy: budgets fail open for a bounded time with a conservative local allowance; allocations fail closed
Leaked reservationsUsage creeps up; customers hit limits they are not usingTTL on every reservation, a sweeper, and reconciliation against the resource inventory
Retries double countA timed-out create is retried and reserves twiceCaller-supplied reservation IDs; commit and release are idempotent
Limit lowered below usageExisting resources exceed the new limitNever delete resources; block new allocations and show usage greater than limit clearly
Hot root counterContention on a large organisation's rowShard the counter with rebalanced sub-limits
Cross-region budgetsEach region admits against a stale global viewSplit the global budget into regional allocations and rebalance them periodically

Whether to fail open or closed is a product decision, not an engineering default. Failing closed on an API budget turns a quota outage into a full outage for every customer; failing open on GPU allocation can oversubscribe hardware. Write the policy per resource into configuration so it is reviewed, and pair budget quotas with load shedding so that protecting the servers never depends on the quota service being up.

Operating it: errors, visibility and increases

Rejections must be self-explanatory. Return a distinct error (HTTP 429 with a Retry-After header for rate and budget quotas, or a gRPC RESOURCE_EXHAUSTED status) whose body names the quota, the scope, the limit, the current usage and the reset time or increase link. An error that just says "quota exceeded" generates a support ticket every time.

Expose usage to customers through an API and a console page, and alert them at 80 and 95 percent. Internally, dashboard rejections per quota, lease refresh latency, reconciliation drift and reservation sweeper activity. Handle increase requests as a workflow: automatic approval up to a threshold based on account age and payment history, human approval above it, and every change written to an audit log with who approved it and why. Keep limits in the quota store, not in service configuration, so an increase never needs a deploy.

Trade-offs to decide explicitly

  • Accuracy versus latency: per-request transactions are exact and slow; leases are fast with a bounded overshoot you choose.
  • Global versus regional: a global store gives one truth and cross-region latency; regional allocations are fast but need rebalancing.
  • Oversubscription: letting child limits exceed the parent improves utilisation but means a child can be refused while below its own limit, which must be explained in the error.
  • Where the check lives: an API gateway can enforce simple per-key budgets; allocation quotas belong in the service that creates the resource, because only it knows when creation succeeded.

What to do next

  1. List every limit your platform enforces and classify each as rate, budget or allocation.
  2. For each, write down the acceptable overshoot and the fail-open or fail-closed policy.
  3. Implement reserve, commit and release with idempotency keys and a TTL sweeper for allocation quotas.
  4. Replace per-request budget checks with leases, and compute the worst-case overshoot from lease size and frontend count.
  5. Emit a usage event for every consumed unit and build a daily reconciliation report with drift per tenant.
  6. Make every rejection name the quota, scope, limit, usage and how to get more.
  7. Load-test the quota service outage: confirm budget quotas degrade gracefully and allocation quotas refuse safely.
Key takeaway: A quota system is three mechanisms under one name. Rate quotas are token buckets at the edge. Budget quotas are leases handed to frontends, with an overshoot bounded by lease size and reconciled against a metering ledger. Allocation quotas are reserve/commit/release transactions with idempotency keys and TTLs. Model tenants as a tree checked leaf to root, keep limits out of deploys, make every rejection explain itself, and decide fail-open or fail-closed per resource before the quota service has its first outage.