Azure AI Foundry is Microsoft's platform for building on hosted models: you deploy a model, call it from code, wrap it in agents and evaluations, and govern all of it from one Azure resource. In late 2025 Microsoft dropped the Azure AI prefix and the documentation now calls it Microsoft Foundry; the older name still appears in portals, scripts and many blog posts, so treat the two as the same product. This article uses Foundry for both.
The hard questions are what it costs at peak, what happens when traffic doubles, and where prompts are processed. Underneath every Foundry deployment there are GPUs, and the single most important decision you make is how you get access to them: shared and billed per token, reserved by the hour, queued for batch, or rented as dedicated virtual machines. This article explains the resource model from first principles, maps every deployment type to the capacity it buys, works through a sizing example, and ends with the failure modes and a checklist you can act on.
The resource and project model
Foundry has a deliberately flat structure. The top level is a Foundry resource, which in Azure Resource Manager is a Microsoft.CognitiveServices/accounts resource of kind AIServices. It shares that provider namespace with Azure OpenAI, Speech, Vision and Language, which is why existing Azure Policy rules and RBAC actions written for Azure OpenAI keep working after an upgrade. Networking, customer-managed keys, model deployments and connections live on this resource.
Inside it sit projects (Microsoft.CognitiveServices/accounts/projects). A project is a development boundary: agents, uploaded files, conversations and evaluation results are scoped to it, and teams reuse the resource's deployments and connections without asking IT to wire anything again. Connected services such as Storage, Key Vault and Azure AI Search remain separate Azure resources with their own network rules. Older hub-based projects are documented as Foundry (classic); start new work on the resource-and-project model.
Access control follows the same split. Control-plane actions such as creating deployments and projects are distinct from data-plane actions such as building agents or uploading files. The built-in roles were recently renamed to Foundry User, Foundry Owner, Foundry Account Owner and Foundry Project Manager (previously Azure AI User and so on); the role IDs did not change. The least-privilege starting point is Foundry User for each developer, and for each project's managed identity, at the resource scope.
Four ways to get GPU capacity
A model deployment binds a model version to a name you call, plus a SKU that decides where and how it runs. For models sold directly by Azure, the serverless option offers three capacity families, and open or custom models can instead run on managed compute.
- Standard (pay per token). Your requests join a shared pool. Quota in tokens per minute is granted per subscription, region, model and deployment type, and you allocate it across your deployments; and you pay only for tokens processed. When the pool is busy you see more latency variance; when you exceed your quota you get HTTP 429. Microsoft recommends starting with Global Standard: it gets new models first, has the lowest price and the highest default quota.
- Provisioned (PTU). You buy provisioned throughput units, a normalised measure of model processing capacity, billed hourly. Capacity is reserved for your deployment, so latency is lower and more predictable. Each model and version needs a different minimum number of PTUs and yields different throughput per PTU, so always size with Microsoft's capacity calculator for the exact model rather than a rule of thumb.
- Batch. You upload a file of requests and get results back with a 24-hour target turnaround at 50 percent of the Global Standard price. Batch has its own enqueued-token quota, so a large job does not starve your online traffic.
- Managed compute. For open-weight or custom models you can deploy onto dedicated virtual machines that are billed while they run, whether or not they receive traffic. The deployment types below do not apply here; you choose instance type and count, and you own the scaling decision.
Deployment types and where prompts are processed
Each serverless family comes in up to three data-processing scopes. Global may process a prompt in any Azure region. Data Zone keeps processing inside a Microsoft-defined zone: the US, the EU Data Boundary, or Asia Pacific. Standard and Regional Provisioned keep processing inside the Azure geography you chose. Data at rest always stays in the resource's geography; the scope only governs where inference runs. The SKU code is what you pass in templates and scripts.
| Deployment type | SKU code | Processing | Billing |
|---|---|---|---|
| Global Standard | GlobalStandard | Any Azure region | Per token |
| Global Provisioned | GlobalProvisionedManaged | Any Azure region | Reserved PTU, hourly |
| Global Batch | GlobalBatch | Any Azure region | 50% of Global Standard, 24 h target |
| Data Zone Standard | DataZoneStandard | Within the data zone | Per token |
| Data Zone Provisioned | DataZoneProvisionedManaged | Within the data zone | Reserved PTU, hourly |
| Data Zone Batch | DataZoneBatch | Within the data zone | Batch pricing |
| Standard | Standard | Within the Azure geography | Per token |
| Regional Provisioned | ProvisionedManaged | Within the Azure geography | Reserved PTU, hourly |
| Developer | DeveloperTier | Any region, no residency guarantee | Per token; fine-tuned model evaluation only, 24 h lifetime, no SLA |
New models arrive in a fixed order: Global first, then Data Zone, then geography-based types, which have no guaranteed date. That ordering is a capacity statement: a global pool can use whichever datacentre has free accelerators, while a geography-scoped deployment can only use GPUs in one geography, which free up as older models retire. If your compliance team requires a single geography, expect to wait for new models and to get lower default quota.
Calling Foundry from code
The current Python SDK is azure-ai-projects 2.x, which targets the new Foundry projects API and is not compatible with 1.x code. You authenticate with Microsoft Entra ID through DefaultAzureCredential, ask the project client for an OpenAI-compatible client, and call the Responses API with your deployment name as the model.
# pip install "azure-ai-projects>=2.3.0" azure-identity
import random, time
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
from openai import RateLimitError, APIStatusError
ENDPOINT = "https://<resource>.services.ai.azure.com/api/projects/<project>"
project = AIProjectClient(endpoint=ENDPOINT, credential=DefaultAzureCredential())
# No SDK-internal retries: a 429 must spill over at once, not after a wait.
client = project.get_openai_client().with_options(max_retries=0)
# Ordered by preference: reserved capacity first, shared pool second.
DEPLOYMENTS = ["support-ptu", "support-global-std"]
def answer(question: str, max_attempts: int = 4) -> str:
for attempt in range(max_attempts):
for name in DEPLOYMENTS:
try:
r = client.responses.create(model=name, input=question)
return r.output_text
except RateLimitError:
continue # this pool is full, try the next one
except APIStatusError as e:
if e.status_code < 500:
raise # 4xx other than 429: a bug, not capacity
# every pool refused: back off with jitter before the next round
time.sleep(min(30, 2 ** attempt) * random.uniform(0.5, 1.0))
raise RuntimeError("all deployments saturated")By default the OpenAI Python client retries a 429 itself and waits on the retry hint the service sends, which would delay the spillover until the reserved pool had already made the user wait. The code therefore turns those retries off with with_options(max_retries=0) and does its own jittered backoff only after every pool has refused. Agents use the same client. You register a versioned agent definition once and then bind a client to it:
from azure.ai.projects.models import PromptAgentDefinition
project.agents.create_version(
agent_name="support-agent",
definition=PromptAgentDefinition(
model="support-ptu",
instructions="Answer from the knowledge base. Say when you do not know.",
),
)
agent_client = project.get_openai_client(agent_name="support-agent")
conv = agent_client.conversations.create()
reply = agent_client.responses.create(conversation=conv.id, input="How do I reset my token?")Deployments themselves belong in infrastructure code rather than the portal, so a review can see the SKU and capacity. With the Azure CLI that is az cognitiveservices account deployment create with --sku-name GlobalStandard (or any SKU code from the table) and --sku-capacity, whose unit is quota for standard SKUs and PTUs for provisioned ones.
Worked example: sizing a support assistant
Suppose you run a customer-support assistant. The numbers here are illustrative; substitute your own measurements. At the daily peak the service handles 25 requests per second, each with about 1,800 input tokens (system prompt, retrieved passages, conversation) and 300 output tokens. Overnight it drops to 3 requests per second. A nightly job also summarises 400,000 closed tickets.
Step 1: convert to tokens per minute. Peak is 25 x 2,100 x 60, about 3.15 million tokens per minute; the overnight floor is about 380,000. Output tokens cost more GPU time than their count suggests because generation is sequential.
Step 2: decide what must be predictable. The assistant is user-facing with a latency target, so the steady daytime load is a good fit for provisioned capacity. Size a Global Provisioned deployment for roughly the daytime baseline, not the peak; the calculator tells you how many PTUs that model needs for that rate. Reserving for the peak means paying for idle accelerators most of the day.
Step 3: absorb the peak with a shared pool. Put a Global Standard deployment of the same model behind it and route 429s from the provisioned deployment to it, exactly as the code above does. You pay per token only for the burst, and you accept more latency variance for that slice of traffic.
Step 4: move the bulk job off the online path. The ticket summaries are not latency-sensitive, so submit them to a Global Batch deployment. It costs half as much per token and uses a separate quota, so it cannot push the daytime traffic into 429s.
Step 5: check residency. If the tickets contain EU customer data and policy requires EU processing, replace every Global SKU with its Data Zone equivalent in an EU region and re-check model availability and quota, both of which are usually lower than Global.
Failure modes
These are the failures that actually page people, and what to do about each.
- 429s at moderate load. Quota belongs to the subscription, region, model and deployment type, and deployments share it. A second deployment in the same region does not add quota unless you request more. Track tokens per minute against quota, not just request counts.
- Regional outage. Foundry does not fail over automatically across regions, and for Global Standard and Data Zone Standard, traffic routed to a region that is having an incident is affected. If you need multi-region availability, deploy separate Foundry resources in two regions and route between them in your application or gateway.
- Latency spikes on shared pools. Standard deployments are best effort; heavy consistent users see more latency variance. If p99 matters, move the baseline to provisioned capacity rather than adding retries.
- Model retirement. Every model version has a retirement date. Pin versions, put retirement dates in your calendar, and run your evaluation set against the replacement before it becomes forced.
- SDK breakage. Code written for
azure-ai-projects1.x does not run on 2.x. Pin the major version in requirements. - Locked-down networking. Some fully private setups cannot be configured from the portal; use the CLI or templates, and give Storage, Key Vault and AI Search their own private endpoints.
- Content filter blocks. Filters run inline per deployment; classify their errors separately so they are not retried as capacity problems.
Operations, governance and cost trade-offs
Run Foundry like any other shared platform. Enable diagnostic settings on the resource to send logs to Log Analytics; resource-level metrics show token consumption, latency, request counts and errors across projects, while agent and evaluation metrics are scoped to each project. Alert on 429 rate and on tokens per minute as a fraction of quota, because the second one warns you before users notice.
Use Azure Policy to stop a team from quietly choosing a deployment type that breaks your residency rules. The rule matches deployments whose sku.name is a forbidden SKU and denies them:
{
"mode": "All",
"policyRule": {
"if": {
"allOf": [
{ "field": "type", "equals": "Microsoft.CognitiveServices/accounts/deployments" },
{ "field": "Microsoft.CognitiveServices/accounts/deployments/sku.name", "equals": "GlobalStandard" }
]
},
"then": { "effect": "deny" }
}
}On cost, the trade-off is the familiar one from any GPU fleet. Per-token pricing is cheapest when utilisation is low or spiky; reserved capacity is cheapest when it stays busy; batch is cheapest of all when nobody is waiting. Managed compute only wins when you need a model or a configuration the serverless catalogue does not offer, because you pay for idle VMs. For how latency behaves on the accelerator underneath, see GPU inference latency and inference optimisation; for why reserved capacity is scarce in the first place, GPU supply gives the background. Azure ML Studio covers the managed-compute side, and cloud AI platforms compares Foundry with its peers.
What to do next
- Create one Foundry resource and one project in a region that offers your target model, and grant Foundry User to developers and to the project's managed identity.
- Deploy the model as Global Standard first and measure real input and output tokens per request for a week.
- Convert peak and baseline to tokens per minute, and check them against your quota and the PTU calculator for that exact model version.
- Decide residency up front: if you need a data zone or geography, switch SKUs now and re-check model availability.
- If latency must be predictable, add a provisioned deployment for the baseline and route its 429s to the standard deployment.
- Move every non-interactive bulk job to a Batch deployment.
- Put deployments in infrastructure code, and add an Azure Policy that denies the SKUs your compliance rules forbid.
- Alert on 429 rate and on quota utilisation, pin model versions and SDK major versions, and record model retirement dates.
- If you need multi-region resilience, stand up a second resource in another region and route between them yourself.