Most teams build the front door of an LLM API like that of any REST service: a gateway, an API key check and a requests-per-minute limit. That is not enough. A single 200-byte request to a model endpoint can ask for tens of thousands of output tokens, carry a system prompt the caller should never control, name a model the caller's plan does not include, attach a PDF with hidden instructions or hold a streaming connection open for minutes. The cost, the risk and the attack surface live in the request body, not in the request count.

Ingress control is the set of checks a request must pass before any model spends a token on it. This page builds that layer from first principles as an ordered admission pipeline, shows the code for its most important stage, token-budget admission, walks through a worked abuse case with numbers and covers the failure modes and trade-offs. It deliberately stays on the inbound side: what leaves the model is covered in egress filtering, and the general gateway architecture (routing, fallbacks, caching) in LLM gateway architecture.

Advertisement

Why an LLM front door is different

Three properties make LLM ingress harder. The first is cost asymmetry. An LLM request costs input plus output tokens times a per-model price, and the caller controls most of it through prompt length, max_tokens, the model name and, for agents, tool rounds. Counting requests treats a one-word question and a 30,000-token essay as equal.

The second is that content is the attack surface. An LLM has no separation between code and data: instructions and data share one token stream. A request can carry a jailbreak, a pasted document with indirect instructions, or fields such as a client-supplied system message or tool definitions that change what the model may do. Ingress is the cheapest place to decide which fields a caller may set at all.

The third is duration. Responses stream for seconds or minutes. A caller who opens a thousand streams and reads them slowly ties up upstream capacity and your provider concurrency quota even if each request looks innocent. Rate limits that count only arrivals miss this completely.

The admission pipeline

Ingress admission pipeline: cheapest checks first, every rejection carries a reason code1. EdgeTLS, size, WAF2. Identitykey / token / mTLS3. Entitlementplan, model, tools4. Validateschema, normalize5. Budgetreserve tokens6. Screeninjection, policy413401403400429block / flagDecision logtenant, key id, check, reason code, estimated tokens, latency; never the raw prompt by defaultRoutermodel + modeModel providerstreamed responseadmittedReconcileactual vs reservedRejections exit downward at the first failing stage; only admitted requests ever cost model tokens.
Six ordered stages. Each one can reject with a specific status and reason code; only admitted requests reach the router.

Order matters: cheap checks first drop garbage before tokenization or a classifier call, and identity comes before everything that is per tenant.

  1. Edge. TLS, a body-size cap derived from the largest context you sell, header limits, per-IP connection rate and a WAF.
  2. Identity. Resolve the caller to a tenant, an application and, where relevant, an end user.
  3. Entitlement. Check the requested model, context length, tools, files and system prompt against the plan.
  4. Validation and normalization. Strict schema, clamped numbers, normalized text, inspected attachments.
  5. Budget admission. Reserve the request's worst-case tokens against the tenant budget and concurrency limit.
  6. Screening. Content checks that block, flag or route to a restricted mode.
Advertisement

Identity and entitlement: authorize the fields, not just the endpoint

A common mistake is to treat authentication as the whole of access control: the key is valid, so the request is allowed. For LLM APIs the interesting authorization questions are about individual fields. Write them down as an entitlement record per tenant and check the request against it.

# entitlements.yaml - one record per plan, overridable per tenant
plans:
  team:
    models: [small-chat, large-chat]
    max_context_tokens: 128000
    max_output_tokens: 8192
    tokens_per_minute: 400000
    max_concurrent_streams: 20
    allow_system_prompt: true       # false on consumer plans: fixed server side
    allow_tools: [search, calculator]
    allow_files: [application/pdf, text/plain]

Three rules make this work. First, never trust client-asserted identity or role fields: a "user_id" in the JSON body is a claim, not an identity; take the tenant and user from the verified credential. Second, for consumer-facing apps, strip or reject client-supplied system messages and assemble the system prompt server side; a client that can write the system message can turn off every behavioural instruction you rely on. Third, store API keys as hashes, scope them (which models, which environments) and give them an owner and an expiry, so one leaked key does not unlock every capability of the tenant.

Validation and normalization

LLM request bodies are nested JSON: messages with roles and content parts, tool definitions, sampling parameters and file references. Validate them with a strict schema that rejects unknown fields; ignored fields are how an unreviewed provider parameter slips through a pass-through gateway.

  • Clamp, do not trust, numeric fields. Set max_tokens to the minimum of the requested value and the plan limit, and reject a missing value rather than defaulting to the model maximum. Bound temperature and the number of completions per request.
  • Bound structure. Cap message count, content parts, part length and the number and size of tool definitions.
  • Normalize text. Apply Unicode NFKC normalization before any screening, and strip or flag zero-width and bidirectional control characters and tag characters, which can hide instructions from human reviewers and naive filters while the model still reads them.
  • Inspect attachments. Sniff the real content type instead of trusting the declared one, enforce page and byte limits, and extract text in a sandbox. Text extracted from a file is untrusted input with the same status as a retrieved web page.
  • Do not fetch URLs at ingress. If the API accepts image or document URLs, fetching them inside the request path is a server-side request forgery risk. Either require uploads, or fetch through an egress proxy with an allowlist and no access to internal address ranges.

Budget admission: reserve tokens, not requests

This is the stage that most front doors are missing. Instead of asking whether the caller has made too many requests, ask whether the caller can afford the most this request could cost. The upper bound is known before the call: the input tokens (which you can count with the model's tokenizer, or estimate conservatively from characters) plus the clamped max_tokens. Reserve that amount against the tenant's token budget, run the request, then refund the difference between the reservation and the tokens actually used.

import math, time

class TokenBudget:
    """Token bucket measured in tokens, with reservation and refund.
    In production this state lives in Redis with an atomic script per tenant."""
    def __init__(self, rate_per_min, burst):
        self.rate = rate_per_min / 60.0
        self.capacity = burst
        self.tokens = burst
        self.updated = time.monotonic()

    def _refill(self):
        now = time.monotonic()
        self.tokens = min(self.capacity, self.tokens + (now - self.updated) * self.rate)
        self.updated = now

    def reserve(self, amount):
        self._refill()
        if amount > self.tokens:
            wait = math.ceil((amount - self.tokens) / self.rate)
            return False, wait            # caller gets 429 with Retry-After: wait
        self.tokens -= amount
        return True, 0

    def refund(self, amount):
        self.tokens = min(self.capacity, self.tokens + amount)


def admit(req, ent, budget, streams, count_tokens):
    max_out = min(req.get("max_tokens") or 0, ent["max_output_tokens"])
    if max_out <= 0:
        return reject(400, "max_tokens_required")
    prompt_tokens = count_tokens(req["messages"])
    if prompt_tokens + max_out > ent["max_context_tokens"]:
        return reject(400, "context_too_large")
    if streams.active(req.tenant) >= ent["max_concurrent_streams"]:
        return reject(429, "concurrency_limit")
    ok, wait = budget.reserve(prompt_tokens + max_out)
    if not ok:
        return reject(429, "token_budget", retry_after=wait)
    return Admitted(reserved=prompt_tokens + max_out, max_out=max_out)

# after the response completes (or is cancelled):
#   budget.refund(admitted.reserved - (usage.input_tokens + usage.output_tokens))

Reservation is conservative, because each in-flight request holds its worst case, and it yields an honest Retry-After. If different models have different prices, reserve in cost units (for example, micro-dollars) instead of raw tokens, using the per-model input and output prices, so one budget covers every model. The broader algorithm choices for distributed limiters, such as GCRA and local leasing, are covered in rate limiter architecture in depth.

Pair the token budget with a concurrency limit per tenant, counted as open streams, and with a daily or monthly spend cap checked at the same point. Token rate protects capacity minute to minute; the spend cap protects the invoice when a leaked key is used slowly for a week.

Worked example: the max_tokens flood

Take a public API on the team plan above: 400,000 tokens per minute, 20 concurrent streams, 8,192 output tokens maximum. An attacker with a stolen key writes a script that sends a five-word prompt asking the model to repeat a word forever, with max_tokens set to 100,000, as fast as possible.

With a limit of 600 requests per minute, every request passes, each runs to the model's own output limit, and the attacker burns millions of output tokens a minute until someone notices the bill.

With the pipeline, the request passes identity (the key is valid) and entitlement (the model is allowed). Validation clamps max_tokens to 8,192. Budget admission reserves about 8,200 tokens per request, so the bucket admits at most about 48 such requests per minute (400,000 divided by 8,200), and the concurrency limit caps simultaneous streams at 20. The tenant's worst-case burn is now bounded by its plan, and the 49th request in a minute gets a 429 with reason token_budget.

Bounded is not detected, though. The decision log shows a key whose requests always ask for the maximum output, use almost all of it and come from a new network: that pattern should feed an anomaly alert and an automatic key suspension rule. Ingress control limits the blast radius; detection, covered in LLM abuse, shortens the window.

Screening at ingress: useful, but not a security boundary

The last stage looks at the content itself: a prompt injection or jailbreak classifier, a check for prohibited use, and detectors for data the caller should not send, such as card numbers in a consumer chat app. Be precise about what this buys. Classifiers are probabilistic, attackers iterate against them, and indirect injection usually arrives later through retrieval or tool results. Ingress screening reduces volume and raises the cost of casual attacks; it does not make a model safe to give dangerous tools.

Use a graded policy. A high-confidence match is blocked with a reason code. A medium-confidence match is admitted in a restricted mode: no tools, no sensitive retrieval, a stricter output filter. A low score is admitted and logged so you can tune thresholds. Give screening a latency budget, and decide in advance whether a timeout fails open or closed for each tenant tier.

The real boundary for agents is downstream: least-privilege tools, per-call authorization and confirmation for destructive actions, as described in LLM tool abuse. Ingress decides which capabilities a request may even ask for; tool authorization decides what each call may do.

Streaming and long-lived connections

Streaming responses need their own limits because the expensive part happens after admission. Set a maximum stream duration and an idle timeout on the client side of the connection, count open streams against the tenant's concurrency limit, and release the slot when the stream ends for any reason.

Most important, propagate cancellation upstream. When a client disconnects, the gateway must cancel the provider request, otherwise you keep paying for tokens nobody reads, and an attacker can open and drop connections to generate cost for free. Test this explicitly: start a long generation, kill the client, and confirm in the provider usage data that generation stopped. Finally, reconcile usage from the provider's reported token counts, not your own estimate, and refund or charge the difference against the reservation.

Operating the front door

Log a structured decision per request: tenant, key identifier, deciding stage, reason code, estimated and actual tokens, screening scores and per-stage latency. Log a hash rather than the raw prompt unless a retention policy says otherwise. Roll out every new check in shadow mode first and compare what it would reject with known-good traffic.

Watch rejections by reason code and tenant, the reserved-to-actual token ratio, screening latency, open streams per tenant and spend per key per day. Alert on a key that departs from its own history rather than on global thresholds.

Failure modes

  • Requests-only rate limiting: a few large-output requests exhaust capacity or budget while every limiter stays green.
  • Pass-through schemas: unknown fields reach the provider, including parameters that enable features the tenant never paid for.
  • Client-controlled system prompt: a consumer app that forwards the client's system message has no stable behaviour to defend.
  • Estimate drift: character-based estimates undercount some languages; reconcile against actual usage.
  • No cancellation: disconnected streams keep generating and billing.
  • Classifier as gate of last resort: a team that relies on ingress screening for agent safety discovers indirect injection through a retrieved document the classifier never saw.
  • Fail-open by accident: the budget store times out and the gateway's error handling admits everything; decide and test the failure policy for every stage.

Trade-offs

ChoiceBenefitCost
Exact tokenizer count at ingressAccurate reservations and context checksCPU per request; needs the right tokenizer per model
Reserve max_tokensHard upper bound on burnAdmits fewer concurrent requests when callers over-ask
Strict schema, reject unknown fieldsNo unreviewed parameters upstreamBreaks clients when providers add fields
Fail closed on budget store outageNo unbounded spendOutage of the store is an outage of the API

What to do next

  1. Write the entitlement record for each plan: models, context and output limits, tools, files, system prompt, token rate, streams and spend cap.
  2. Switch your schema validation to reject unknown fields and clamp max_tokens to the plan limit; reject requests without it.
  3. Implement token reservation with refund, and a per-tenant open-stream limit, in front of every model route.
  4. Verify that client disconnects cancel upstream generation, using provider usage data as the proof.
  5. Add reason codes to every rejection and a structured decision log without raw prompts.
  6. Run any content screening in shadow mode for a week, then enforce with a block, restrict or log policy.
  7. Add a per-key anomaly alert on output tokens per hour and an automatic suspension path.
Key takeaway: LLM ingress control is an ordered admission pipeline: edge limits, verified identity, field-level entitlements, strict validation and normalization, token-budget reservation and graded content screening, cheapest first and each with a reason code. Count and reserve tokens or cost rather than requests, never let public clients set the system prompt or unknown fields, cancel upstream when a stream drops, and treat ingress screening as volume reduction rather than the boundary that makes tools safe.