OpenAI's Moderation API is a free classifier that takes text, images or both and returns, for each of thirteen harm categories, a score between 0 and 1 and a boolean flag. It is the cheapest safety signal most teams can add to an LLM product: one HTTP call, no model to host and no bill. That ease is also what makes it easy to misuse. Teams wire the top-level flagged field straight to a block decision, send it categories it cannot see in images, assume it detects prompt injection, and then discover months later that the model behind the -latest alias moved and their false-positive rate doubled.

This article covers the API as it is documented today: the request and response shapes, the category matrix, the newer option of requesting moderation inside a generation call, how to turn scores into your own policy, a worked calibration, the production wiring around timeouts and failures, and where the API stops being the right tool. For the general architecture of a moderation system, layered classifiers, review queues and appeals, read LLM moderation architecture first; this page is the concrete guide to one component of it.

What the endpoint is, and is not

The endpoint is a multi-label classifier trained against OpenAI's own content taxonomy. Multi-label matters: a single input can score high on harassment and violence at once, and each category is judged independently. The current model family is omni-moderation, which accepts text and images and does not classify audio. The API reference lists omni-moderation-latest, an alias that OpenAI upgrades over time, and a dated snapshot, omni-moderation-2024-09-26. The older text-only models (text-moderation-007, text-moderation-stable and text-moderation-latest) were shut down on 27 October 2025 according to OpenAI's deprecations page, so code that still names them is broken, not merely old.

What it is not is just as important. It is not a prompt-injection or jailbreak detector: an instruction to ignore the system prompt contains no harmful content and scores near zero. It does not know your product policy, so a medical support tool and a children's tutoring app get identical scores for the same sentence. It is not a CSAM detection system; the documentation says explicitly not to send known or suspected child sexual abuse material to it and to use dedicated child-safety tooling instead. And it is not a fact checker, a PII scanner or a data-leak filter. Treat it as one well-calibrated sensor among several, as in the layered design of defence in depth for LLM applications.

Request and response anatomy

A request carries a model and an input. The input can be a string, an array of strings (one result per string) or an array of typed parts mixing text and image_url objects, where the URL can be an https link or a base64 data URL. Images can be up to 20 MB. The response holds an id, the model that served it and a results array, each element with four fields: flagged, categories (per-category booleans), category_scores (per-category floats) and category_applied_input_types, which lists for each category whether the score came from text, image or both. Category keys contain slashes, such as self-harm/intent, so the raw JSON is the clearest thing to log. The following uses plain HTTP so the keys are exactly what the API returns:

import os, httpx

def moderate(text=None, image_url=None, model="omni-moderation-2024-09-26"):
    parts = []
    if text:
        parts.append({"type": "text", "text": text})
    if image_url:
        parts.append({"type": "image_url", "image_url": {"url": image_url}})
    r = httpx.post(
        "https://api.openai.com/v1/moderations",
        headers={"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"},
        json={"model": model, "input": parts},
        timeout=3.0,
    )
    r.raise_for_status()
    body = r.json()
    res = body["results"][0]
    return body["model"], res["category_scores"], res["category_applied_input_types"]

The categories, and which ones can be scored from an image, are fixed by the model. This matrix decides what an image-only upload can and cannot tell you:

CategoryInputs scoredNote
harassment, harassment/threateningText onlyHate against non-protected groups is classed as harassment
hate, hate/threateningText onlyProtected attributes such as race, religion, disability, caste
illicit, illicit/violentText onlyAdvice or instructions for illicit acts; the violent variant adds weapons or violence
self-harm, self-harm/intent, self-harm/instructionsText and imagesIntent means the speaker is doing or planning it
sexualText and imagesExcludes sex education and wellness
sexual/minorsText onlyNot an image CSAM detector; see above
violence, violence/graphicText and imagesGraphic means depicted in graphic detail

If you send only an image, the text-only categories come back as 0. That zero means the category was not evaluated, not that the image is clean, which is why category_applied_input_types belongs in every log line.

Standalone versus inline moderation

There are now two ways to get scores. The standalone endpoint classifies anything you send it, with no generation involved: user messages, uploaded images, retrieved documents, or text from another model. Alternatively, a Responses or Chat Completions request can carry a top-level moderation object naming the moderation model, and the API returns scores for the input and for the generated output with the response, with no second round trip. The model still generates normally; nothing is blocked for you. Your code reads the results and decides.

Two ways to call moderation, and where your policy sitsUser inputtext and/or imagePOST /v1/moderationsstandalone, freePolicy engineyour thresholdsBlock / reviewrefuse, queue, logscoresdenyallowGeneration requestmoderation: {model: omni-moderation-latest}Response + moderation.input + moderation.outputeach is a result or an error; streamed output is scored at the endOutput policysame thresholds, output sideAudit log: model id, per-category scores, applied input types, decision, policy versionthe raw material for recalibration when the model behind -latest changes
Standalone moderation gates input before generation; inline moderation returns input and output scores with the response. Both feed the same policy engine and audit log.
from openai import OpenAI
client = OpenAI()

resp = client.responses.create(
    model=GEN_MODEL,                      # your generation model
    input=[{"role": "user", "content": user_text}],
    moderation={"model": "omni-moderation-latest"},
)
for side in ("input", "output"):
    result = getattr(resp.moderation, side)
    if result.type == "error":            # moderation could not complete
        raise ModerationUnavailable(side, result.message)
    raw = result.to_dict()                # API field names, slashes included
    action, reasons = decide(raw["category_scores"], raw["category_applied_input_types"])

The decide function is the policy table defined in the next section. Four documented details shape how you use this. Each side can hold an error instead of scores, so check the type before reading. With Chat Completions and several choices, output result i belongs to choice i. For tool-calling requests, moderation covers tool-call arguments and tool outputs that appear in the conversation, but not tool names, descriptions or schemas. And when you stream, the scores arrive only after the full output exists, not with each delta. That last point decides the architecture: inline moderation cannot stop a harmful sentence that has already been streamed to the screen. If the output must be checked before the user sees it, buffer the output, or run the standalone endpoint on chunks as described in content safety, in depth.

From scores to your own policy

The flagged boolean reflects OpenAI's default thresholds, chosen for a general audience. The documentation itself says to treat scores as signals for your own policy rather than automatic blocking decisions, and notes that even a refusal that discusses harmful content can be flagged. A support bot for a crisis line needs to let users talk about self-harm and route them to help; a game chat needs to block harassment at a much lower score than a news summariser. So the policy should be a table you own, versioned like code, with an action per category and per side:

POLICY_VERSION = "2026-10-07.1"
RULES = {
    # category:           (block_at, review_at)   input side
    "self-harm/intent":       (None, 0.30),   # never block: route to help flow
    "self-harm/instructions": (0.50, 0.20),
    "harassment/threatening": (0.60, 0.25),
    "hate/threatening":       (0.50, 0.20),
    "illicit/violent":        (0.60, 0.30),
    "sexual/minors":          (0.10, 0.02),   # most conservative
    "violence/graphic":       (0.80, 0.50),
}

def decide(scores, applied):
    action, reasons = "allow", []
    for cat, (block_at, review_at) in RULES.items():
        s = scores.get(cat, 0.0)
        if not applied.get(cat):               # not evaluated for this input type
            continue
        if block_at is not None and s >= block_at:
            return "block", [(cat, s)]
        if review_at is not None and s >= review_at:
            action, reasons = "review", reasons + [(cat, s)]
    return action, reasons

The thresholds above are illustrative starting points, not recommendations. The only thresholds worth shipping are ones you measured on your own traffic, which is what the next section does. Note the skip for categories not evaluated: with that check, an image-only post is never cleared by a score that was never computed.

Worked example: calibrating a threat threshold

Suppose you run a community forum and want to block threatening harassment in replies. You sample 2,000 recent replies, oversampling reports so the positive class is not vanishingly rare, and have two reviewers label each against your written definition. Say 120 are true threats. You run every reply through the pinned snapshot and sweep the harassment/threatening threshold. These numbers are illustrative; yours will differ, which is the point:

ThresholdThreats caught (of 120)RecallBenign blocked (of 1,880)Precision
0.1011293%14144%
0.2510487%4769%
0.409378%1983%
0.607764%693%

There is no free point on the curve. At 0.60 you block few innocent replies but miss a third of threats; at 0.10 almost half of what you block is benign. A two-tier policy resolves it: block automatically at 0.60, where precision is high, and send the 0.25 to 0.60 band to human review. Under these numbers, reviewers see about 27 real threats and 41 benign replies per 2,000, a workload you can staff. Recompute per language and per surface (titles, replies, direct messages), because score distributions shift with context. Then freeze the table, tag it with the snapshot name, and rerun the same labelled set whenever you consider moving to a new model. A drop of more than a few points in recall at your chosen threshold is a migration blocker, not a footnote.

Production wiring

Moderation sits on the hot path, so its failure behaviour is a product decision. Decide per surface whether to fail open (allow, log, re-check asynchronously) or fail closed (hold the message with a retry prompt). Public posting surfaces usually fail closed; a private assistant answering its own user often fails open with after-the-fact review.

  • Timeouts. Give the call a short, explicit timeout and one retry with jitter on 5xx or 429 responses. Do not retry indefinitely in a request thread; queue the re-check instead.
  • Batching. When classifying stored content such as backfills or retrieved documents, send arrays of strings so one request returns one result per string, and keep requests well under your account's rate limits; the guide does not publish numbers, so read them from your organisation's limits page and response headers.
  • Long text. Classify long documents in chunks at paragraph boundaries and take the maximum score per category. A single harmful paragraph inside a long benign document is otherwise diluted, and you lose the location.
  • Caching. Cache results keyed on a hash of the exact input and the model name. Never reuse a cached score after changing snapshot.
  • Logging. Record the returned model, all scores, applied input types, the decision and POLICY_VERSION. Store hashes or redacted text where retention rules require it.
  • Monitoring. Chart the daily distribution of each category's score, not just the block rate. A shift in the median with flat traffic means the model or the population changed.

Failure modes

  • Alias drift. omni-moderation-latest is upgraded in place, and the docs warn that policies on category_scores may need recalibration. Pin the dated snapshot in production and move to a new one deliberately.
  • Image blind spots. Harassment, hate, illicit and sexual/minors are text-only. A hateful meme with no caption scores 0 on hate. Extract text from images with OCR and send it as a text part, or add an image classifier for those categories.
  • Dead model names. Code calling the text-moderation models has failed since October 2025. Search your repositories and configuration for them.
  • Streaming leak. Inline scores arrive after the full output, so a streamed reply is already on screen when the flag lands.
  • Context loss. Each request is scored on what you send. A threat built over five messages, each harmless alone, needs a window of recent turns in the input.
  • Language skew. Calibrate separately for each major language on your platform; a threshold tuned on English traffic is not evidence for any other language.
  • Category mismatch. Your policy may care about things the taxonomy lacks, such as medical misinformation, spam or competitor promotion. Those need their own classifiers.
  • Wrong tool. Using it as an injection or jailbreak defence leaves that hole open; see jailbreak defence for controls that address it.

Trade-offs

Against a self-hosted safety model such as Llama Guard, the Moderation API costs nothing per call and needs no GPU, but your content leaves your infrastructure, the taxonomy is fixed, and the model changes on OpenAI's schedule unless you pin. Llama Guard lets you write custom categories into the prompt and keep data in-house, at the cost of hosting and latency. Against a classifier trained on your own labels, the API is weaker on your specific harms but needs no labelling pipeline to start. A common end state uses all three: the API as a broad, cheap first screen, a self-hosted model for categories you define, and a small custom classifier for the one or two harms that matter most to your product.

What to do next

  1. Search your code for text-moderation model names and migrate any you find to the omni model.
  2. Pin omni-moderation-2024-09-26 (or whichever snapshot you validate) in production and record the model name on every decision.
  3. Write your policy table: category, side, block threshold, review threshold, owner.
  4. Label 1,000 to 2,000 items of your own traffic and sweep thresholds per category and per language.
  5. Add OCR text extraction for image uploads so the text-only categories get evaluated.
  6. Decide fail-open or fail-closed per surface and test it by blocking the endpoint in staging.
  7. If you stream, decide whether to buffer output or accept post-hoc retraction, and document the choice.
  8. Chart daily score distributions per category and alert on shifts.
Key takeaway: The Moderation API is a free, multi-label, text-and-image classifier with thirteen fixed categories. Use its scores, not its flag, as input to a policy table you calibrate on your own traffic; pin a dated snapshot; log applied input types so text-only categories are never silently cleared on images; and decide explicitly what happens when the call fails or the output has already streamed.