Azure's prebuilt AI APIs let you call speech recognition, translation, OCR, document extraction, text analytics and content moderation over HTTPS without training a model. They have been renamed several times, which makes old tutorials confusing, but the engineering underneath has been stable: you create an account resource in a region, it gets an endpoint, and every request carries either a key or a Microsoft Entra ID token.

This article explains that model from first principles. It untangles the names, shows how the resource is laid out, walks through the service catalogue and what is being retired, and builds a worked document-processing pipeline with keyless authentication. It then covers private networking, throttling and retries, running the models in containers, the common ways these integrations fail, and when a prebuilt API is a better choice than prompting a large language model.

Advertisement

What the name means now

The services launched as Azure Cognitive Services. In 2023 Microsoft unified them as Azure AI services, and in late 2025, with the introduction of Microsoft Foundry, the documentation began calling them Foundry Tools and the multi-service account a Foundry resource. The current authentication page is titled "Authentication in Foundry Tools". Existing endpoints, SDKs and APIs kept working through each rename.

What did not change is the Azure Resource Manager type. Every one of these accounts is a Microsoft.CognitiveServices/accounts resource, distinguished by its kind. Single-service kinds such as TextAnalytics, SpeechServices or FormRecognizer expose one API. The older multi-service kind is CognitiveServices. The current multi-service kind is AIServices, which is also what Foundry builds projects on. When you read infrastructure code, the kind tells you what the resource can do far more reliably than the portal's display name.

How the resource is laid out

There are two planes. The control plane is Azure Resource Manager: you create the account, choose a pricing tier, list and rotate keys, and set network rules there, and access is governed by Azure RBAC roles on the resource. The data plane is the account's endpoint: the actual API calls for sentiment, transcription or OCR go there.

The endpoint comes in two forms. A regional endpoint such as westus.api.cognitive.microsoft.com is shared by everyone in the region. A custom subdomain such as https://contoso-ai.cognitiveservices.azure.com/ is unique to your resource. The custom subdomain is not cosmetic: Microsoft's documentation states that Entra authentication always needs it and regional endpoints do not support Entra tokens, and private endpoints depend on it as well. Set a custom subdomain on every resource you create.

One resource, two planes: ARM manages the account, the custom subdomain serves the APIsYour applicationmanaged identityMicrosoft Entra IDtoken for cognitiveservicesAzure Resource Managercreate, keys, network rulesAccount (kind AIServices)Custom subdomainname.cognitiveservices...Network rulesprivate endpoint, firewallAuthEntra token or keyQuota and meteringper service, per regionLanguageSpeechVisionDocument IntelligenceTranslator, Content Safety1. get token2. HTTPS callcontrol planeKeys bypass Entra entirely: disable local auth once every caller uses tokens.
The application gets an Entra token and calls the resource's custom subdomain; Resource Manager handles keys, network rules and configuration. Each service has its own quota and metering.
Advertisement

The service catalogue

ServiceWhat it doesNotes
LanguageSentiment, key phrases, named entities, PII detection and redaction, summarization, conversational language understanding (CLU), custom question answeringCLU and custom question answering succeed LUIS and QnA Maker
SpeechSpeech to text, text to speech, speech translation, batch transcriptionSupports short-lived access tokens as well as keys
VisionImage analysis, OCR (Read), spatial analysisComputer Vision API v1.0 to v3.1 retired on September 13, 2026
Document IntelligenceLayout, prebuilt invoice, receipt and ID models, custom extraction modelsFormerly Form Recognizer; the v2.0 API retired on August 31, 2026
TranslatorText and document translationGlobal endpoint; a multi-service key also needs the region header
Content SafetySeverity scores for hate, sexual, violence and self-harm in text and images; prompt shieldsOften placed in front of and behind LLM calls
Anomaly Detector, Metrics Advisor, PersonalizerTime-series anomaly detection, metrics monitoring, recommendationsRetired on October 1, 2026; no new resources since September 2023

Azure OpenAI and other Foundry models also live under the AIServices kind but are a separate subject with their own deployment and quota model. Retirement dates above come from Microsoft's 2026 lifecycle page; check it for anything you depend on, because API versions retire independently of services.

Provisioning with the CLI

RG=rg-ai-prod; LOC=westeurope; NAME=contoso-ai-prod

az cognitiveservices account create \
  --name $NAME --resource-group $RG --location $LOC \
  --kind AIServices --sku S0 \
  --custom-domain $NAME \
  --assign-identity --yes

# Grant your application's managed identity data-plane access.
APP_PRINCIPAL=$(az webapp identity show -g $RG -n contoso-api --query principalId -o tsv)
SCOPE=$(az cognitiveservices account show -g $RG -n $NAME --query id -o tsv)
az role assignment create --assignee $APP_PRINCIPAL \
  --role "Cognitive Services User" --scope $SCOPE

# Once every caller uses Entra tokens, turn keys off.
az resource update --ids $SCOPE --set properties.disableLocalAuth=true

Role assignments can take several minutes to propagate, so a 401 immediately after provisioning is not necessarily a mistake. The free tier, F0, is useful for experiments but has low rate limits and is not available for every service; S0 is the usual production tier.

Authentication: keys, tokens and Entra ID

There are three mechanisms. A resource key goes in the Ocp-Apim-Subscription-Key header. Each resource has two keys so you can rotate one while the other is in use. A multi-service key used against Translator must also send Ocp-Apim-Subscription-Region. Some services, including Translator and Speech, accept a short-lived access token obtained by posting a key to the regional /sts/v1.0/issueToken endpoint; those tokens are valid for 10 minutes, which is useful for handing a browser or device a credential that expires quickly.

The third mechanism, and the one to standardize on, is Microsoft Entra ID. Your code obtains a token for the scope https://cognitiveservices.azure.com/.default using a managed identity, workload identity or developer login, and sends it as Authorization: Bearer. The caller needs a data-plane role such as Cognitive Services User on the resource. No secret is stored anywhere, access appears in sign-in logs, and revoking a role takes effect without rotating anything. Once every client uses tokens, setting disableLocalAuth closes the key path entirely. Background on identities and conditional access is in the Entra ID deep dive.

# pip install azure-identity azure-ai-textanalytics azure-ai-contentsafety
from azure.identity import DefaultAzureCredential
from azure.ai.textanalytics import TextAnalyticsClient
from azure.ai.contentsafety import ContentSafetyClient
from azure.ai.contentsafety.models import AnalyzeTextOptions

ENDPOINT = "https://contoso-ai-prod.cognitiveservices.azure.com/"
cred = DefaultAzureCredential()           # managed identity in Azure, az login locally

language = TextAnalyticsClient(ENDPOINT, cred)
safety = ContentSafetyClient(ENDPOINT, cred)

docs = ["Call Maria on 555-0100 about invoice 4471."]
for result in language.recognize_pii_entities(docs):
    if not result.is_error:
        print(result.redacted_text)        # PII replaced with asterisks

verdict = safety.analyze_text(AnalyzeTextOptions(text=docs[0]))
for item in verdict.categories_analysis:
    print(item.category, item.severity)

Worked example: an invoice pipeline

Suppose finance wants supplier invoices that arrive as PDFs turned into structured records, with personal data redacted before anything reaches the analytics store. The pipeline has four steps: the PDF lands in a storage container (see Data Lake Storage Gen2), a worker sends it to Document Intelligence's prebuilt invoice model, the worker redacts PII in free-text fields with Language, and the cleaned record is written out.

Document Intelligence analysis is a long-running operation. The submit call returns 202 Accepted with an Operation-Location header, and you poll that URL until the status is succeeded or failed. The SDK hides the polling, but the REST shape is worth knowing because it is the same pattern used by batch transcription and document translation. The request below targets the v4.0 GA API version; check the current version in the REST reference before you deploy.

import time, requests
from azure.identity import DefaultAzureCredential

ENDPOINT = "https://contoso-ai-prod.cognitiveservices.azure.com"
token = DefaultAzureCredential().get_token("https://cognitiveservices.azure.com/.default").token
H = {"Authorization": f"Bearer {token}"}

def analyze_invoice(pdf_url: str) -> dict:
    submit = requests.post(
        f"{ENDPOINT}/documentintelligence/documentModels/prebuilt-invoice:analyze",
        params={"api-version": "2024-11-30"},
        headers=H, json={"urlSource": pdf_url}, timeout=30)
    submit.raise_for_status()                       # expect 202; a 429 here should be retried too
    poll_url = submit.headers["Operation-Location"]
    while True:
        r = requests.get(poll_url, headers=H, timeout=30)
        if r.status_code == 429:                    # throttled: honour Retry-After
            time.sleep(int(r.headers.get("Retry-After", "2"))); continue
        body = r.json()
        if body["status"] in ("succeeded", "failed"):
            return body
        time.sleep(2)

doc = analyze_invoice("https://contosodocs.blob.core.windows.net/in/inv-4471.pdf")
fields = doc["analyzeResult"]["documents"][0]["fields"]
print(fields.get("VendorName", {}).get("content"), fields.get("InvoiceTotal", {}).get("content"))

Two details matter in production. Each extracted field carries a confidence score, so route low-confidence invoices to a human review queue rather than writing them straight through. And when the worker passes a blob URL, the service fetches it itself, so the blob must be reachable by the service; with private storage, send the bytes in the request body instead.

Networking

By default the endpoint is public and protected only by authentication. For production, add network rules. The account firewall can allow selected virtual networks and IP ranges. A private endpoint goes further: it places a private IP for the resource inside your virtual network, and with public access disabled, traffic never leaves Microsoft's backbone. The pattern is the same as for any PaaS service, described in Private Link.

Private endpoints need DNS. The custom subdomain must resolve to the private IP inside your network, which is done with a private DNS zone, privatelink.cognitiveservices.azure.com for the cognitive services endpoint. An AIServices account can expose additional hostnames, such as the OpenAI-compatible one, that need their own zones; take the exact list from the private endpoint page for your resource rather than guessing. The classic failure is a private endpoint that works from one network and times out from another because that network resolves the public address.

Quotas, throttling and retries

Each service enforces its own rate limits per resource, measured as transactions per second or per minute, and they differ by tier. Exceeding them returns HTTP 429, usually with a Retry-After header. The Azure SDKs retry 429 and transient 5xx responses with exponential backoff by default; raw HTTP clients must do it themselves, as the polling loop above does.

Design for limits rather than around them. Put a queue in front of batch workloads so a backlog becomes latency instead of errors. Use the batch APIs where they exist, such as batch transcription and document translation, instead of firing thousands of synchronous calls. Watch the resource's metrics for total calls, throttled calls and latency, and alert on the throttled-call rate. If a single resource is not enough, split by workload across resources, possibly in different regions, which also gives you a failover target.

Containers and disconnected use

Several models, including parts of Language, Speech, Translator, Read OCR and Document Intelligence, are published as Docker containers. A container runs the model on your hardware, which helps with data residency, latency and bursty local workloads. Standard containers still need network access to Azure to report usage for billing; they are started with the endpoint and key of a resource in your subscription.

docker run --rm -p 5000:5000 --memory 8g --cpus 4 \
  mcr.microsoft.com/azure-cognitive-services/textanalytics/language:latest \
  Eula=accept Billing=https://contoso-ai-prod.cognitiveservices.azure.com/ ApiKey=$KEY

Fully disconnected containers exist for environments with no internet connection, but they require an approved application and a commitment-tier plan. Take the image name and tag for each model from its container documentation; tags and resource requirements differ by model.

Failure modes

  • 401 with Entra tokens. The resource has no custom subdomain, the token was requested for the wrong scope, or the role assignment has not propagated yet.
  • Keys leaked in code. A key grants full data-plane access to every service on a multi-service account. Use tokens, keep any remaining keys in Key Vault, and disable local auth.
  • Retired API versions. Code pinned to an old API version starts failing on its retirement date. Inventory the api-version parameters and SDK versions you call.
  • Region mismatch. Features and models are not available in every region. Check availability before choosing where to deploy, especially for data-residency requirements.
  • Silent quality drift. Prebuilt models are updated by Microsoft. Pin a model version where the API allows it and keep a labelled test set to rerun after upgrades.
  • Throttling storms. Clients that retry immediately and in lockstep turn a brief limit into a long outage. Use jittered backoff and a queue.

Prebuilt APIs versus prompting an LLM

Many of these tasks can now be done by prompting a general model, so choose deliberately. Prebuilt APIs return a fixed schema with confidence scores, are cheap per call, have predictable latency and are easy to evaluate. They are the better choice for high-volume, well-defined tasks such as OCR, invoice fields, PII redaction and speech transcription. An LLM is better when the task is open-ended, the schema changes often or reasoning across a whole document is needed. A common design uses both: prebuilt extraction and redaction first, then an LLM over the cleaned output, with Content Safety on either side. For a wider comparison across clouds, see cloud AI platforms.

What to do next

  1. Inventory every Cognitive Services resource and record its kind, region, tier, custom subdomain and whether local auth is enabled.
  2. Check each workload against the 2026 retirements and API-version retirements, and plan migrations for anything affected.
  3. Move callers to managed identity with the Cognitive Services User role, then set disableLocalAuth.
  4. Add a private endpoint and private DNS zone for production resources, and test name resolution from every network that calls them.
  5. Put a queue and jittered retries in front of batch workloads, and alert on throttled calls.
  6. Build a small labelled test set for each prebuilt model you rely on and rerun it after every model or API upgrade.
Key takeaway: Azure Cognitive Services, Azure AI services and Foundry Tools are three names for the same thing: Microsoft.CognitiveServices accounts, now usually of kind AIServices, exposing prebuilt speech, language, vision, document and safety APIs. Give every resource a custom subdomain, call it with Entra ID tokens from managed identities, then disable keys and put it behind a private endpoint with correct DNS. Respect per-service rate limits with queues and backoff, track model and API-version retirements, including the services retired on October 1, 2026, and use prebuilt APIs for well-defined high-volume tasks, with LLMs where open-ended reasoning is needed.