Azure OpenAI Service puts OpenAI models behind an Azure resource: your subscription, your region, your network rules, your identity provider and your quota. The model weights and the chat or responses API look almost the same as OpenAI's own platform, so it is tempting to treat the service as a URL swap. Most production incidents come from the parts that differ: how a request is admitted against a tokens-per-minute budget before the model has produced a single token, how content filtering shows up as either an HTTP 400 or a successful response with a truncated answer, and how a 429 can mean four different things.
This article follows one request from your code to the model and back. It covers the v1 endpoint and keyless authentication, the arithmetic of admission control, the rate-limit headers and what to do with each kind of 429, content-filter handling, network and identity hardening, and a gateway that splits one deployment's quota between teams. The choice of deployment type (Global, Data Zone, regional, provisioned, batch) and spillover between provisioned and standard capacity are covered in Azure AI Foundry and GPU capacity; here they get one paragraph so we can stay on the request path.
Resources, deployments and quota
Three nouns matter. A resource is the Azure object with an endpoint such as https://contoso-ai.openai.azure.com, a region, network settings and access control. A deployment lives inside a resource and binds a model and version to a deployment type and a capacity. A model is what the deployment runs. In API calls the model field carries the deployment name, not the model name, which lets you upgrade a deployment to a new model version without changing code and also means a typo in the deployment name is a 404, not a silent fallback.
Quota is assigned per subscription, per region, per model and per deployment type, in tokens per minute (TPM). When you create a deployment you carve TPM out of that pool. A subscription with 240,000 TPM of one model in one region could run one deployment at 240K, two at 120K, or any split that sums to the total. A requests-per-minute (RPM) limit is derived from the TPM through capacity units whose ratio depends on the model: Microsoft's quota page lists 6 RPM per 1,000 TPM for older chat models and very different ratios for reasoning models, so read the table for your model rather than assuming one ratio. Provisioned deployments are sized in provisioned throughput units (PTUs) and use a utilization-based admission check instead of the token counter described below.
The path of one request
The diagram shows the stages. The client obtains a Microsoft Entra ID token (or presents a key), optionally passes through a gateway, and reaches the resource endpoint, where network rules decide whether the caller may connect at all. The deployment's admission control then estimates the request's token cost and checks it against the per-minute counters. The prompt passes through the content-filter classifiers, the model generates, and the completion is filtered again before it is returned with rate-limit headers.
Each stage fails differently, and the application has to tell them apart: a 401 or 403 is identity or role assignment, a connection refused or 403 with a network message is the firewall or private endpoint, a 404 is the deployment name, a 429 is admission, a 400 with code content_filter is the prompt filter, and a normal 200 whose finish_reason is content_filter is the output filter.
Calling the v1 API without keys
The v1 API removes the old requirement to pin a dated api-version query parameter and lets the standard OpenAI client talk to Azure directly. Point base_url at the resource endpoint with /openai/v1/ appended and pass a token provider as api_key; the client calls the provider before each request, so tokens refresh without your code noticing. Current Microsoft samples request the scope https://ai.azure.com/.default; older samples use https://cognitiveservices.azure.com/.default. Grant the calling identity the Cognitive Services OpenAI User role on the resource.
# pip install openai azure-identity
from openai import OpenAI
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
token_provider = get_bearer_token_provider(
DefaultAzureCredential(), "https://ai.azure.com/.default"
)
client = OpenAI(
base_url="https://contoso-ai.openai.azure.com/openai/v1/",
api_key=token_provider, # a callable, refreshed automatically
max_retries=0, # we handle 429s ourselves, see below
)
raw = client.chat.completions.with_raw_response.create(
model="support-chat", # the DEPLOYMENT name
messages=[{"role": "user", "content": "Summarise ticket 4411 in two lines."}],
max_tokens=200,
)
resp = raw.parse()
print(resp.choices[0].message.content)
print(raw.headers.get("x-ratelimit-remaining-tokens"),
raw.headers.get("x-ratelimit-remaining-requests"))Two choices in that snippet are deliberate. DefaultAzureCredential resolves to a managed identity in Azure and to your developer login locally, so no key ever sits in configuration. And with_raw_response exposes the response headers, which are the only accurate view of how close the deployment is to its limit; the usage metrics in Azure Monitor lag and count billed tokens, not admitted ones.
How admission control counts tokens
Admission control is the part that surprises people. When a request arrives, the service computes an estimated maximum token count from the prompt text, the max_tokens setting and best_of, adds it to a per-minute counter, and returns 429 for the rest of the minute once the counter passes the deployment's TPM. The estimate is partly character-based and is not the billed count, which is computed after generation. Rejected requests still count. Separately, the RPM limit is enforced over short windows, typically 1 or 10 seconds, so a 600 RPM deployment can throttle a burst of eleven requests in one second while being far under 600 for the minute.
Worked example. A support summariser runs on a deployment with 100,000 TPM. Prompts average 1,500 tokens and answers average 250 tokens, but a developer set max_tokens to 4,000 "to be safe". Each request is admitted at roughly 1,500 + 4,000 = 5,500 tokens, so the deployment admits about 100,000 / 5,500, or 18 requests per minute, while Azure Monitor shows only 18 x 1,750 = 31,500 tokens used. The team sees 429s at a third of their quota and files for an increase. Setting max_tokens to 400 drops the estimate to about 1,900 tokens and raises admitted throughput to about 52 requests per minute with no quota change. Trimming a 600-token boilerplate system prompt raises it to about 77.
Microsoft's quota page names max_tokens explicitly. Newer models take max_completion_tokens in chat completions or max_output_tokens in the responses API, and the page does not spell out how those enter the estimate. Treat them as output caps that probably count the same way, and confirm by reading x-ratelimit-remaining-tokens before and after a single request on an idle deployment.
Reading 429s and rate-limit headers
Every response carries x-ratelimit-limit-requests, x-ratelimit-limit-tokens, the matching remaining and reset headers, and a 429 carries retry-after-ms. Microsoft documents four causes of 429, and each needs a different response:
| Cause | How to recognise it | Right response |
|---|---|---|
| TPM or RPM exceeded | "Rate limit is exceeded"; remaining-tokens near zero | Back off; shrink output caps; rebalance or request quota |
| Shared capacity pressure | "temporarily unable to process" or "high demand" | Retry after retry-after-ms; consider provisioned capacity |
| Temporary limit adjustment | x-ratelimit-limit-tokens lower than your configured TPM | Retry with backoff; spread traffic; it usually clears in hours |
| Inflated estimate | 429s while billed-token metrics look low | Lower max_tokens and best_of; trim prompts |
A client that classifies its own 429s and paces itself from the headers:
import random, time
from openai import RateLimitError
def call_with_pacing(client, configured_tpm, **kwargs):
for attempt in range(6):
try:
raw = client.chat.completions.with_raw_response.create(**kwargs)
h = raw.headers
remaining = int(h.get("x-ratelimit-remaining-tokens", "1000000"))
limit = int(h.get("x-ratelimit-limit-tokens", str(configured_tpm)))
if limit < configured_tpm:
log_metric("aoai.limit_adjusted", limit) # temporary adjustment in force
if remaining < 0.1 * limit:
time.sleep(1.0) # slow down before the wall
return raw.parse()
except RateLimitError as e:
ms = e.response.headers.get("retry-after-ms")
wait = int(ms) / 1000 if ms else min(30, 2 ** attempt)
time.sleep(wait * random.uniform(0.8, 1.2)) # jitter: no thundering herd
raise RuntimeError("deployment still throttling after 6 attempts")The SDK retries 429s twice by default and honours the retry hint; that is fine for a script but wrong if you also wrap calls in your own retry loop, because the attempts multiply. Pick one layer. The spillover pattern, where a 429 from a provisioned deployment is sent immediately to a standard one, needs SDK retries off for the same reason and is shown in the Foundry article linked above.
Content filtering in application code
Content filtering runs classifiers over the prompt and over the completion. When the prompt is blocked the request fails with HTTP 400 and error code content_filter; no tokens are generated. When the completion is blocked the request succeeds with HTTP 200, but the choice's finish_reason is content_filter and the content may be missing or cut off. Responses also carry prompt_filter_results and per-choice content_filter_results annotations naming the categories and severities. In streaming mode the filter signal can arrive after text has been sent; Microsoft documents that it arrives within about 1,000 characters of the offending content and that both prompt and completion tokens are billed when it fires.
from openai import BadRequestError
def safe_answer(client, messages):
try:
resp = client.chat.completions.create(
model="support-chat", messages=messages, max_tokens=400)
except BadRequestError as e:
if e.code == "content_filter":
audit("prompt_blocked", e.body) # categories are in the error body
return "I can't help with that request."
raise # other 400s are bugs: fail loudly
choice = resp.choices[0]
if choice.finish_reason == "content_filter":
audit("completion_blocked", resp.to_dict().get("prompt_filter_results"))
return "The answer was withheld by policy. Please rephrase."
if choice.finish_reason == "length":
audit("truncated", resp.usage.completion_tokens)
return choice.message.contentThe common bug is treating every 200 as a complete answer. A chat UI that streams tokens should also be able to retract text already shown when the stream ends with a filter reason.
Identity and network hardening
Production resources should accept only Entra tokens and only private traffic. Set the resource's disableLocalAuth property to true so API keys stop working; any key that leaked into a notebook becomes useless. Give each workload its own managed identity with the Cognitive Services OpenAI User role, which can call deployments but cannot create or change them. Put the resource behind a private endpoint and disable public network access; the custom subdomain then resolves to a private IP through a private DNS zone, as described in private connectivity to cloud services. Where a key cannot be avoided, for example a third-party tool, keep it in Azure Key Vault and rotate it on a schedule using the two keys each resource has. The identity design for agents that call models on a user's behalf is covered in workload identity for LLM systems.
Sharing a deployment through a gateway
One deployment often serves several teams, and Azure's quota is per deployment, so one noisy team can starve the others. Azure API Management solves this with the azure-openai-token-limit policy, which keeps a token counter per key of your choosing and returns 429 to the caller that exceeds its share before the request reaches the model. With estimate-prompt-tokens enabled, the gateway counts prompt tokens up front, so a team already over budget never consumes backend admission.
<policies>
<inbound>
<base />
<authentication-managed-identity resource="https://cognitiveservices.azure.com" />
<azure-openai-token-limit
counter-key="@(context.Subscription.Id)"
tokens-per-minute="20000"
estimate-prompt-tokens="true" />
</inbound>
</policies>Here each APIM subscription, issued one per team, gets 20,000 TPM of a larger deployment, and the gateway authenticates to Azure OpenAI with its own managed identity so teams never hold backend credentials. Give the per-team budgets a sum below the deployment's TPM, because the backend still applies its own estimate, and log the gateway's token metrics per team for chargeback.
Failure modes
Failures seen repeatedly in real deployments:
- 429s at a fraction of quota. Inflated
max_tokensor bursts inside a 1-second RPM window. Fix the cap and smooth the request rate with a client-side token bucket. - Retry storms. SDK retries inside an application retry loop inside a gateway retry policy; one user request becomes dozens. Retry at one layer only.
- Silent truncation.
finish_reasonoflengthorcontent_filtertreated as success. Check it on every response and alert on the rate. - Model retirement. Model versions have retirement dates; a deployment pinned to a retired version stops serving. Track the dates and test the next version on a parallel deployment before you switch.
- Primary region outage. Even Global Standard and Data Zone Standard traffic first routed to the resource's region is affected; put a second resource in another region behind the gateway.
- Key leakage. Keys pasted into scripts outlive their authors. Disable local auth.
Operating it and the trade-offs
Watch three signals per deployment: the 429 rate split by cause (from your own logs, since only the client sees the message text), the ratio of billed tokens to admitted tokens (a low ratio means inflated output caps), and the finish_reason distribution. Review quota allocation monthly, because quota parked on an idle deployment is unavailable to a busy one in the same region. Keep prompts and completions out of application logs unless your data-handling policy allows it, and follow the broader checklist in LLM deployment hardening.
The trade-off: you gain private networking, Entra identity and regional processing, and in exchange manage quota yourself and treat filtering as part of your application's behaviour.
What to do next
- List every deployment with its model version, deployment type and TPM, and note each model's retirement date.
- Switch clients to the v1 endpoint with a token provider, then set
disableLocalAuthon the resource. - Grep your code for
max_tokensand set each cap from the 99th percentile of real completion lengths. - Log every 429 with its message and the rate-limit headers, and classify it into the four causes.
- Keep exactly one retry layer and make it honour
retry-after-ms. - Handle both content-filter paths: the 400 with
content_filterand the 200 withfinish_reasoncontent_filter. - If several teams share a deployment, add an APIM token-limit policy with per-team budgets.
- Move the resource behind a private endpoint and disable public network access.