Azure AI Content Safety is Microsoft's hosted moderation service. You send it text or an image and it returns a severity per harm category; you send it a prompt and some documents and it tells you whether either looks like an attack on a language model; you send it a model answer plus the sources it was meant to use and it tells you which parts are not supported by them. Called directly, it is a plain REST API you can put in front of any model, including ones not hosted on Azure.
The service does not decide anything for you. It returns numbers and booleans; your code turns them into allow, block, rewrite or escalate. Most production problems come from that gap: a threshold copied from a quickstart, a document longer than the endpoint accepts, or an outage that silently turned moderation off. This article covers the endpoints, severity policy, request placement, a working gateway, threshold calibration and operations.
What the service is made of
The service is a set of separate endpoints with separate limits, languages, regions and maturity levels. Treat each as its own dependency. The figures below were taken from Microsoft Learn in October 2026; API versions are deprecated 90 days after a compatible successor ships, so recheck them when you pin a version.
| Capability | Endpoint | Returns | Notes |
|---|---|---|---|
| Analyze text | text:analyze | Severity for Hate, Sexual, Violence, SelfHarm; blocklist matches | GA; 10K characters per call; harm models trained on eight languages |
| Analyze image | image:analyze | Severity for the same four categories | GA; up to 4 MB, 50x50 to 7200x7200 pixels |
| Prompt Shields | text:shieldPrompt | attackDetected for the prompt and for each document | GA; prompt up to 10K characters; up to five documents totalling 10K |
| Groundedness detection | text:detectGroundedness | Ungrounded flag, ungrounded proportion, spans | Preview; English; few regions |
| Protected material | text:detectProtectedMaterial | Whether output matches known text such as lyrics or articles | GA; English; minimum 110 characters |
| Blocklists | text/blocklists/{name} | Exact-term matches inside analyze text | 10,000 terms across all lists; 128 characters per term |
Two boundaries matter. First, the service states that it cannot be used to detect child sexual abuse imagery; that needs a dedicated hash-matching service. Second, guardrails configured on a Foundry model deployment and direct calls to a Content Safety resource are different products with different limits; for example, Foundry guardrails moderate the first 1,000 characters of text scenarios, while the direct analyze text API accepts 10,000. Know which one you are relying on.
Severity levels and turning them into policy
Analyze text classifies each input on four categories independently, so one message can be both violent and hateful. With "outputType": "FourSeverityLevels" (the default) each category gets 0, 2, 4 or 6, meaning safe, low, medium and high. With "EightSeverityLevels" you get 0 to 7, which gives finer control at the boundary that usually matters, between low and medium.
A severity is an ordinal label, not a calibrated probability. Severity 4 does not mean a 4-in-6 chance of harm; it means the model placed the content in its medium band of the published severity definitions, and the right threshold depends on your surface, not on the scale. A children's education app may block Violence at 2; a news-summarisation tool that handles war reporting may allow Violence up to 4 on output and block at 6.
Write the policy down as a table keyed by surface, direction and category, and load it from configuration so it can change without a deploy:
# policy.yaml: thresholds are "block at or above"
support_chat:
input: {Hate: 4, Sexual: 4, Violence: 4, SelfHarm: 2}
output: {Hate: 2, Sexual: 2, Violence: 4, SelfHarm: 2}
selfharm_route: crisis_resources # SelfHarm >= 2 on input changes the reply, not just blocks itNote the asymmetry: input thresholds are usually looser than output thresholds, because users quote, complain and describe harm they have suffered, and blocking their message is a worse failure than declining to repeat it. Self-harm deserves its own action. A user describing self-harm should get a supportive response with resources, not a refusal.
Where the calls sit in a request
In a retrieval-augmented assistant there are three places to call the service and one optional fourth. The input gate runs analyze text and Prompt Shields on the user message, in parallel because they are independent. The document gate runs Prompt Shields on retrieved chunks, which is where indirect injection lives: a web page or email that tells the model to exfiltrate data. The output gate runs analyze text and, for long generations, protected material detection on the answer. Groundedness detection can run on the answer and its sources, but it is preview, slower and region-limited, so most teams run it on a sample, asynchronously, as a quality signal rather than a gate.
Two design rules keep this sane. Every gate writes one decision record holding the scores, the action, the policy version and the API version, but not the raw text unless your data policy allows it. And every gate has an explicit failure mode for when the service times out: fail closed on high-risk surfaces, fail open with a flag on low-risk ones. Decide this per surface in advance; do not let an exception handler decide it by accident.
A moderation gateway in Python
The gateway below uses the Python SDK for analyze text, as in Microsoft's quickstart, and plain REST for Prompt Shields with the documented field names. It chunks long text to stay under the 10,000-character limit, takes the maximum severity across chunks, and returns a decision rather than raising.
import os, requests
from concurrent.futures import ThreadPoolExecutor
from azure.ai.contentsafety import ContentSafetyClient
from azure.ai.contentsafety.models import AnalyzeTextOptions
from azure.core.credentials import AzureKeyCredential
from azure.core.exceptions import AzureError # covers HTTP, timeout and connection errors
ENDPOINT = os.environ["CONTENT_SAFETY_ENDPOINT"].rstrip("/")
KEY = os.environ["CONTENT_SAFETY_KEY"] # prefer Entra ID / managed identity in production
client = ContentSafetyClient(ENDPOINT, AzureKeyCredential(KEY))
API = "2024-09-01"
MAX_CHARS = 10_000
def chunks(text, size=MAX_CHARS - 200, overlap=200):
step = size - overlap
return [text[i:i + size] for i in range(0, max(len(text), 1), step)]
def severities(text):
worst = {}
for part in chunks(text):
res = client.analyze_text(AnalyzeTextOptions(text=part))
for item in res.categories_analysis:
name = getattr(item.category, "value", item.category) # enum -> "Hate"
worst[name] = max(worst.get(name, 0), item.severity or 0)
return worst
def shield(user_prompt, documents):
attack, doc_flags = False, []
for i, part in enumerate(chunks(user_prompt)): # never truncate: scan every chunk
r = requests.post(
f"{ENDPOINT}/contentsafety/text:shieldPrompt",
params={"api-version": API},
headers={"Ocp-Apim-Subscription-Key": KEY},
json={"userPrompt": part, "documents": documents if i == 0 else []},
timeout=3)
r.raise_for_status()
body = r.json()
attack |= body["userPromptAnalysis"]["attackDetected"]
doc_flags = doc_flags or [d["attackDetected"] for d in body.get("documentsAnalysis", [])]
return attack, doc_flags
def input_gate(message, policy, fail_closed=True):
try:
with ThreadPoolExecutor(2) as pool:
sev_f = pool.submit(severities, message)
sh_f = pool.submit(shield, message, [])
sev, (attack, _) = sev_f.result(), sh_f.result()
except (AzureError, requests.RequestException) as exc:
return {"action": "block" if fail_closed else "allow_flagged", "error": str(exc)}
if attack:
return {"action": "block", "reason": "prompt_attack", "scores": sev}
hits = {k: v for k, v in sev.items() if v >= policy.get(k, 7)}
return {"action": "block" if hits else "allow", "hits": hits, "scores": sev}Retrieved documents need batching too: one call accepts up to five documents with a combined 10,000 characters, so a top-8 retrieval of 2,000-character chunks needs at least two calls. Drop or quarantine any chunk flagged as an attack instead of discarding the whole answer; a single poisoned page should not take down every query that retrieves it.
Prompt Shields: a boolean, not a score
Prompt Shields returns a boolean per input, not a score. That is simpler to use and harder to tune: there is no threshold to move when it over-blocks a legitimate power-user prompt. If false positives hurt, your levers are around it. Route a flagged prompt to a stricter mode (no tools, no retrieval) instead of refusing outright.
It is also one layer, not a solution. Classifiers trained on known attack styles miss new ones. The structural defences still apply: least-privilege tools, confirmation before side-effects, and never giving retrieved text the authority of the system prompt.
Groundedness and protected material
Groundedness detection takes the model output (text, up to 7,500 characters), the groundingSources (up to 55,000 characters in total), a task of QnA or Summarization and a domain of GENERIC or MEDICAL. It returns ungroundedDetected, an ungroundedPercentage and the ungrounded spans. Microsoft's docs are explicit that the percentage is the proportion of text judged ungrounded, not a confidence. Optional reasoning and correction modes call an Azure OpenAI deployment you supply, adding latency and cost. It is in preview, English-only and available in a handful of regions, so build it as a sampled quality metric before you consider it as a gate.
Protected material detection is for outputs: it flags generations that reproduce known text such as song lyrics or articles, and it ignores inputs shorter than 110 characters. Run it on long answers, not on every short reply.
Blocklists
Blocklists add exact terms that the classifiers will not know: internal codenames, a competitor's leaked product name, slurs specific to your community. Create a list with a PATCH to contentsafety/text/blocklists/{name}, add items in batches of up to 100, and pass blocklistNames on analyze text calls. Setting haltOnBlocklistHit to true skips the classifiers when a term matches, which saves time but loses the severity record for that message.
from azure.ai.contentsafety import BlocklistClient
from azure.ai.contentsafety.models import AddOrUpdateTextBlocklistItemsOptions, TextBlocklistItem
bl = BlocklistClient(ENDPOINT, AzureKeyCredential(KEY))
terms = [t.strip() for t in open("blocklist.txt", encoding="utf-8") if t.strip()]
for i in range(0, len(terms), 100): # 100 items per request; 128 chars each
bl.add_or_update_blocklist_items(
blocklist_name="support-chat-terms",
options=AddOrUpdateTextBlocklistItemsOptions(
blocklist_items=[TextBlocklistItem(text=t) for t in terms[i:i + 100]]))Edits take effect after a delay, documented as usually not more than five minutes, so an emergency term added during an incident is not live instantly; keep a local deny check in the gateway for the first minutes. The 10,000-term cap across all lists means blocklists are for precise terms, not for encoding policy.
Worked example: calibrating output thresholds
Suppose you run a support assistant and start, as many teams do, by blocking output at medium (4) in every category. Before launch you label 600 real, anonymised assistant outputs: 540 acceptable and 60 that your policy team says must not be shown. You run them through analyze text with eight-level output and sweep the Violence and Hate thresholds. The numbers below are an illustration of the method, not published benchmarks.
| Block at | Unsafe caught (of 60) | Acceptable blocked (of 540) | Recall | False block rate |
|---|---|---|---|---|
| 5 | 38 | 2 | 63% | 0.4% |
| 4 | 47 | 6 | 78% | 1.1% |
| 3 | 55 | 19 | 92% | 3.5% |
| 2 | 58 | 61 | 97% | 11.3% |
The starting point of 4 misses 13 of the 60 unsafe outputs. Moving to 3 catches 8 more at the cost of 13 more false blocks in 540. With eight levels you can choose 3; with four you would have to jump to 2 and accept an 11% false block rate. Read the 19 false blocks at threshold 3: if most are refund disputes with angry wording, add an allowlisted rewrite path instead of lowering the threshold. With only 60 positives, the recall estimate has a wide interval (roughly plus or minus 7 points at 92%), so grow the labelled set before you treat the difference between 3 and 4 as settled.
Running it in production
- Rate limits. F0 allows 5 requests per second; S0 allows 1,000 requests per 10 seconds for the moderation and Prompt Shields APIs, and 50 per second for groundedness. One chat turn can use four or five calls, so size from calls per turn, not turns. Back off on 429 with jitter.
- Latency. Run independent calls concurrently, set a tight timeout per call, and measure p95 added latency per gate. For streamed output, moderate in windows (for example every few hundred characters plus the final text) rather than per token.
- Regions and residency. Create the resource where you want data processed. Groundedness is available in only a few regions, and some features use global routing; check the region page before you promise residency.
- Languages. Harm categories were trained and tested on eight languages; Prompt Shields, protected material and groundedness were tested on English. Measure recall per language you serve.
Failure modes
- Silent fail-open. A broad exception handler turns a timeout into allow. Alert on the rate of gate errors, not just on blocks.
- Truncation. Sending only the first 10,000 characters lets harmful text hide at the end of a long document. Chunk, and check the last chunk.
- Encoding evasion. Base64, leetspeak or text split across turns can slip past per-message classifiers. Moderate the decoded or normalised form where you decode, and moderate the conversation window, not just the latest message.
- Threshold drift. New features shift the input distribution; re-sweep quarterly.
Trade-offs and alternatives
Compared with self-hosted classifiers such as Llama Guard, the service gives you managed models, image support and no GPU to run, in exchange for a network hop, per-call cost, per-region availability and a fixed taxonomy you cannot retrain. Compared with AWS Bedrock Guardrails, it is model-agnostic by design, since you call it yourself, but you also own the orchestration Bedrock does inline. For the general layering, see content safety, in depth; for the self-hosted alternative, Llama Guard; for the AWS equivalent, Bedrock Guardrails; and for how to evaluate injection detectors honestly, prompt injection scanners.
What to do next
- List every surface and direction (input, retrieved documents, output) and write the policy table with a threshold and an action per category.
- Build the gateway with chunking, concurrency, timeouts and an explicit fail-open or fail-closed choice per surface.
- Label at least a few hundred real samples per surface and sweep thresholds with eight-level output before launch.
- Run Prompt Shields on retrieved documents, not only on user prompts.
- Log one decision record per gate with scores, action, policy version and API version.
- Load-test against the rate limit at your calls-per-turn ratio and alert on gate error rates.
- Pilot groundedness detection on sampled traffic as a quality metric before making it a gate.
- Schedule a quarterly re-sweep and an API version review.