Azure's prebuilt AI APIs let you call speech recognition, translation, OCR, document extraction, text analytics and content moderation over HTTPS without training a model. They have been renamed several times, which makes old tutorials confusing, but the engineering underneath has been stable: you create an account resource in a region, it gets an endpoint, and every request carries either a key or a Microsoft Entra ID token.
This article explains that model from first principles. It untangles the names, shows how the resource is laid out, walks through the service catalogue and what is being retired, and builds a worked document-processing pipeline with keyless authentication. It then covers private networking, throttling and retries, running the models in containers, the common ways these integrations fail, and when a prebuilt API is a better choice than prompting a large language model.
What the name means now
The services launched as Azure Cognitive Services. In 2023 Microsoft unified them as Azure AI services, and in late 2025, with the introduction of Microsoft Foundry, the documentation began calling them Foundry Tools and the multi-service account a Foundry resource. The current authentication page is titled "Authentication in Foundry Tools". Existing endpoints, SDKs and APIs kept working through each rename.
What did not change is the Azure Resource Manager type. Every one of these accounts is a Microsoft.CognitiveServices/accounts resource, distinguished by its kind. Single-service kinds such as TextAnalytics, SpeechServices or FormRecognizer expose one API. The older multi-service kind is CognitiveServices. The current multi-service kind is AIServices, which is also what Foundry builds projects on. When you read infrastructure code, the kind tells you what the resource can do far more reliably than the portal's display name.
How the resource is laid out
There are two planes. The control plane is Azure Resource Manager: you create the account, choose a pricing tier, list and rotate keys, and set network rules there, and access is governed by Azure RBAC roles on the resource. The data plane is the account's endpoint: the actual API calls for sentiment, transcription or OCR go there.
The endpoint comes in two forms. A regional endpoint such as westus.api.cognitive.microsoft.com is shared by everyone in the region. A custom subdomain such as https://contoso-ai.cognitiveservices.azure.com/ is unique to your resource. The custom subdomain is not cosmetic: Microsoft's documentation states that Entra authentication always needs it and regional endpoints do not support Entra tokens, and private endpoints depend on it as well. Set a custom subdomain on every resource you create.
The service catalogue
| Service | What it does | Notes |
|---|---|---|
| Language | Sentiment, key phrases, named entities, PII detection and redaction, summarization, conversational language understanding (CLU), custom question answering | CLU and custom question answering succeed LUIS and QnA Maker |
| Speech | Speech to text, text to speech, speech translation, batch transcription | Supports short-lived access tokens as well as keys |
| Vision | Image analysis, OCR (Read), spatial analysis | Computer Vision API v1.0 to v3.1 retired on September 13, 2026 |
| Document Intelligence | Layout, prebuilt invoice, receipt and ID models, custom extraction models | Formerly Form Recognizer; the v2.0 API retired on August 31, 2026 |
| Translator | Text and document translation | Global endpoint; a multi-service key also needs the region header |
| Content Safety | Severity scores for hate, sexual, violence and self-harm in text and images; prompt shields | Often placed in front of and behind LLM calls |
| Anomaly Detector, Metrics Advisor, Personalizer | Time-series anomaly detection, metrics monitoring, recommendations | Retired on October 1, 2026; no new resources since September 2023 |
Azure OpenAI and other Foundry models also live under the AIServices kind but are a separate subject with their own deployment and quota model. Retirement dates above come from Microsoft's 2026 lifecycle page; check it for anything you depend on, because API versions retire independently of services.
Provisioning with the CLI
RG=rg-ai-prod; LOC=westeurope; NAME=contoso-ai-prod
az cognitiveservices account create \
--name $NAME --resource-group $RG --location $LOC \
--kind AIServices --sku S0 \
--custom-domain $NAME \
--assign-identity --yes
# Grant your application's managed identity data-plane access.
APP_PRINCIPAL=$(az webapp identity show -g $RG -n contoso-api --query principalId -o tsv)
SCOPE=$(az cognitiveservices account show -g $RG -n $NAME --query id -o tsv)
az role assignment create --assignee $APP_PRINCIPAL \
--role "Cognitive Services User" --scope $SCOPE
# Once every caller uses Entra tokens, turn keys off.
az resource update --ids $SCOPE --set properties.disableLocalAuth=trueRole assignments can take several minutes to propagate, so a 401 immediately after provisioning is not necessarily a mistake. The free tier, F0, is useful for experiments but has low rate limits and is not available for every service; S0 is the usual production tier.
Authentication: keys, tokens and Entra ID
There are three mechanisms. A resource key goes in the Ocp-Apim-Subscription-Key header. Each resource has two keys so you can rotate one while the other is in use. A multi-service key used against Translator must also send Ocp-Apim-Subscription-Region. Some services, including Translator and Speech, accept a short-lived access token obtained by posting a key to the regional /sts/v1.0/issueToken endpoint; those tokens are valid for 10 minutes, which is useful for handing a browser or device a credential that expires quickly.
The third mechanism, and the one to standardize on, is Microsoft Entra ID. Your code obtains a token for the scope https://cognitiveservices.azure.com/.default using a managed identity, workload identity or developer login, and sends it as Authorization: Bearer. The caller needs a data-plane role such as Cognitive Services User on the resource. No secret is stored anywhere, access appears in sign-in logs, and revoking a role takes effect without rotating anything. Once every client uses tokens, setting disableLocalAuth closes the key path entirely. Background on identities and conditional access is in the Entra ID deep dive.
# pip install azure-identity azure-ai-textanalytics azure-ai-contentsafety
from azure.identity import DefaultAzureCredential
from azure.ai.textanalytics import TextAnalyticsClient
from azure.ai.contentsafety import ContentSafetyClient
from azure.ai.contentsafety.models import AnalyzeTextOptions
ENDPOINT = "https://contoso-ai-prod.cognitiveservices.azure.com/"
cred = DefaultAzureCredential() # managed identity in Azure, az login locally
language = TextAnalyticsClient(ENDPOINT, cred)
safety = ContentSafetyClient(ENDPOINT, cred)
docs = ["Call Maria on 555-0100 about invoice 4471."]
for result in language.recognize_pii_entities(docs):
if not result.is_error:
print(result.redacted_text) # PII replaced with asterisks
verdict = safety.analyze_text(AnalyzeTextOptions(text=docs[0]))
for item in verdict.categories_analysis:
print(item.category, item.severity)
Worked example: an invoice pipeline
Suppose finance wants supplier invoices that arrive as PDFs turned into structured records, with personal data redacted before anything reaches the analytics store. The pipeline has four steps: the PDF lands in a storage container (see Data Lake Storage Gen2), a worker sends it to Document Intelligence's prebuilt invoice model, the worker redacts PII in free-text fields with Language, and the cleaned record is written out.
Document Intelligence analysis is a long-running operation. The submit call returns 202 Accepted with an Operation-Location header, and you poll that URL until the status is succeeded or failed. The SDK hides the polling, but the REST shape is worth knowing because it is the same pattern used by batch transcription and document translation. The request below targets the v4.0 GA API version; check the current version in the REST reference before you deploy.
import time, requests
from azure.identity import DefaultAzureCredential
ENDPOINT = "https://contoso-ai-prod.cognitiveservices.azure.com"
token = DefaultAzureCredential().get_token("https://cognitiveservices.azure.com/.default").token
H = {"Authorization": f"Bearer {token}"}
def analyze_invoice(pdf_url: str) -> dict:
submit = requests.post(
f"{ENDPOINT}/documentintelligence/documentModels/prebuilt-invoice:analyze",
params={"api-version": "2024-11-30"},
headers=H, json={"urlSource": pdf_url}, timeout=30)
submit.raise_for_status() # expect 202; a 429 here should be retried too
poll_url = submit.headers["Operation-Location"]
while True:
r = requests.get(poll_url, headers=H, timeout=30)
if r.status_code == 429: # throttled: honour Retry-After
time.sleep(int(r.headers.get("Retry-After", "2"))); continue
body = r.json()
if body["status"] in ("succeeded", "failed"):
return body
time.sleep(2)
doc = analyze_invoice("https://contosodocs.blob.core.windows.net/in/inv-4471.pdf")
fields = doc["analyzeResult"]["documents"][0]["fields"]
print(fields.get("VendorName", {}).get("content"), fields.get("InvoiceTotal", {}).get("content"))Two details matter in production. Each extracted field carries a confidence score, so route low-confidence invoices to a human review queue rather than writing them straight through. And when the worker passes a blob URL, the service fetches it itself, so the blob must be reachable by the service; with private storage, send the bytes in the request body instead.
Networking
By default the endpoint is public and protected only by authentication. For production, add network rules. The account firewall can allow selected virtual networks and IP ranges. A private endpoint goes further: it places a private IP for the resource inside your virtual network, and with public access disabled, traffic never leaves Microsoft's backbone. The pattern is the same as for any PaaS service, described in Private Link.
Private endpoints need DNS. The custom subdomain must resolve to the private IP inside your network, which is done with a private DNS zone, privatelink.cognitiveservices.azure.com for the cognitive services endpoint. An AIServices account can expose additional hostnames, such as the OpenAI-compatible one, that need their own zones; take the exact list from the private endpoint page for your resource rather than guessing. The classic failure is a private endpoint that works from one network and times out from another because that network resolves the public address.
Quotas, throttling and retries
Each service enforces its own rate limits per resource, measured as transactions per second or per minute, and they differ by tier. Exceeding them returns HTTP 429, usually with a Retry-After header. The Azure SDKs retry 429 and transient 5xx responses with exponential backoff by default; raw HTTP clients must do it themselves, as the polling loop above does.
Design for limits rather than around them. Put a queue in front of batch workloads so a backlog becomes latency instead of errors. Use the batch APIs where they exist, such as batch transcription and document translation, instead of firing thousands of synchronous calls. Watch the resource's metrics for total calls, throttled calls and latency, and alert on the throttled-call rate. If a single resource is not enough, split by workload across resources, possibly in different regions, which also gives you a failover target.
Containers and disconnected use
Several models, including parts of Language, Speech, Translator, Read OCR and Document Intelligence, are published as Docker containers. A container runs the model on your hardware, which helps with data residency, latency and bursty local workloads. Standard containers still need network access to Azure to report usage for billing; they are started with the endpoint and key of a resource in your subscription.
docker run --rm -p 5000:5000 --memory 8g --cpus 4 \
mcr.microsoft.com/azure-cognitive-services/textanalytics/language:latest \
Eula=accept Billing=https://contoso-ai-prod.cognitiveservices.azure.com/ ApiKey=$KEYFully disconnected containers exist for environments with no internet connection, but they require an approved application and a commitment-tier plan. Take the image name and tag for each model from its container documentation; tags and resource requirements differ by model.
Failure modes
- 401 with Entra tokens. The resource has no custom subdomain, the token was requested for the wrong scope, or the role assignment has not propagated yet.
- Keys leaked in code. A key grants full data-plane access to every service on a multi-service account. Use tokens, keep any remaining keys in Key Vault, and disable local auth.
- Retired API versions. Code pinned to an old API version starts failing on its retirement date. Inventory the
api-versionparameters and SDK versions you call. - Region mismatch. Features and models are not available in every region. Check availability before choosing where to deploy, especially for data-residency requirements.
- Silent quality drift. Prebuilt models are updated by Microsoft. Pin a model version where the API allows it and keep a labelled test set to rerun after upgrades.
- Throttling storms. Clients that retry immediately and in lockstep turn a brief limit into a long outage. Use jittered backoff and a queue.
Prebuilt APIs versus prompting an LLM
Many of these tasks can now be done by prompting a general model, so choose deliberately. Prebuilt APIs return a fixed schema with confidence scores, are cheap per call, have predictable latency and are easy to evaluate. They are the better choice for high-volume, well-defined tasks such as OCR, invoice fields, PII redaction and speech transcription. An LLM is better when the task is open-ended, the schema changes often or reasoning across a whole document is needed. A common design uses both: prebuilt extraction and redaction first, then an LLM over the cleaned output, with Content Safety on either side. For a wider comparison across clouds, see cloud AI platforms.
What to do next
- Inventory every Cognitive Services resource and record its kind, region, tier, custom subdomain and whether local auth is enabled.
- Check each workload against the 2026 retirements and API-version retirements, and plan migrations for anything affected.
- Move callers to managed identity with the Cognitive Services User role, then set disableLocalAuth.
- Add a private endpoint and private DNS zone for production resources, and test name resolution from every network that calls them.
- Put a queue and jittered retries in front of batch workloads, and alert on throttled calls.
- Build a small labelled test set for each prebuilt model you rely on and rerun it after every model or API upgrade.