If you build on the Gemini API, safety filtering is not one thing you switch on. Three separate mechanisms can stop or change an answer, they report themselves in different places in the response, and one of them does not report itself at all. Teams that treat the filter as a single boolean ship apps that crash on an empty response.text, show users half an answer that was cut mid-stream, or quietly loosen a threshold until something embarrassing gets through.

This article explains the filter from the integrator's side: what each layer does, which settings you control, how to read every block signal, how to pick thresholds from your own data rather than by feel, and which risks you still own after the filter has done its job. Specifics were checked against Google's safety-settings guide and API reference on 2026-10-01; the API evolves, so verify enum names against the reference before you depend on them.

Advertisement

Three layers, not one switch

The first layer is a set of built-in protections against core harms. Google's guide gives content that endangers child safety as the example and says these are always blocked and cannot be adjusted. You cannot configure them, and your app should treat them as a hard edge.

The second layer is the adjustable classifier. Every prompt and every candidate response is scored per harm category, and you choose, per category, how high the score must be before the request or response is blocked. This is what people usually mean by "safety settings".

The third layer is the model itself. Gemini is trained to decline some requests. A trained refusal is not a block: the response arrives normally, with text such as "I can't help with that" and a finish reason of STOP. No setting turns this off, and none of the block fields mention it. If you count only flagged blocks, you undercount how often users fail to get an answer.

Outside all three sits your application layer: your own input checks, output handling and domain policy. The filter knows nothing about your product, your users or your data, so that layer is never optional.

Where a Gemini request can be stopped, and where each stop shows up in the responseYour appinput checks, policyPrompt filterscore + per-categoryModeltrained refusalsResponse filtersper-category, SPII...requestallowedtokenspromptFeedbackblockReason, no textCandidate textmay be a polite refusalfinishReasonSAFETY, BLOCKLIST...blockedanswerblockedYour outcome classifierone function maps every signal to: answered, refused, prompt_blocked, response_blocked, partial_streamThree different things can stop an answer; only two of them are flagged, and a refusal arrives as ordinary text with finishReason STOP
Prompt filters, the model and response filters each stop requests differently. Blocks are flagged in promptFeedback or finishReason; refusals are ordinary text. One classifier function should turn every case into a named outcome.

The knobs you actually have

The guide lists four adjustable categories: HARM_CATEGORY_HARASSMENT, HARM_CATEGORY_HATE_SPEECH, HARM_CATEGORY_SEXUALLY_EXPLICIT and HARM_CATEGORY_DANGEROUS_CONTENT. The API reference also names further categories, such as civic integrity and jailbreak; whether a given model honours them varies, so test before you rely on one.

Each category takes a HarmBlockThreshold. The classifier reports a probability level of NEGLIGIBLE, LOW, MEDIUM or HIGH, and the threshold says which levels are blocked:

ThresholdBlocks when probability isTypical use
BLOCK_LOW_AND_ABOVELOW, MEDIUM or HIGHChildren's products, strict brand surfaces
BLOCK_MEDIUM_AND_ABOVEMEDIUM or HIGHGeneral consumer chat
BLOCK_ONLY_HIGHHIGHProfessional tools where false blocks are costly
BLOCK_NONEnever; content shown regardless of probabilityMeasurement runs, vetted internal tools
OFFfilter turned off for the categorySame as above, where you own the downstream check
HARM_BLOCK_THRESHOLD_UNSPECIFIEDuses the defaultAvoid; be explicit

The most important fact in the table is the default. The guide (checked 2026-10-01) says that when no threshold is set, the default block threshold is Off for Gemini 2.5 and 3 models. If you never pass safety_settings, the adjustable layer is not blocking anything on those models, and you are relying on the core protections, the model's own refusals and whatever you built. Many teams assume the opposite. Set every category explicitly so the behaviour is visible in code review and does not shift when you change model.

from google import genai
from google.genai import types

MODEL = "your-gemini-model-id"   # pick a current id from the models page; do not hard-code a stale one
client = genai.Client()           # reads GEMINI_API_KEY from the environment

SAFETY = [
    types.SafetySetting(category=types.HarmCategory.HARM_CATEGORY_HARASSMENT,
                        threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE),
    types.SafetySetting(category=types.HarmCategory.HARM_CATEGORY_HATE_SPEECH,
                        threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE),
    types.SafetySetting(category=types.HarmCategory.HARM_CATEGORY_SEXUALLY_EXPLICIT,
                        threshold=types.HarmBlockThreshold.BLOCK_LOW_AND_ABOVE),
    types.SafetySetting(category=types.HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT,
                        threshold=types.HarmBlockThreshold.BLOCK_ONLY_HIGH),
]

resp = client.models.generate_content(
    model=MODEL,
    contents=user_text,
    config=types.GenerateContentConfig(safety_settings=SAFETY),
)
Advertisement

Where the verdict shows up

A blocked prompt and a blocked response look different. When the prompt is blocked, prompt_feedback.block_reason is set and there are no candidates to read. The reference lists SAFETY, OTHER, BLOCKLIST, PROHIBITED_CONTENT and IMAGE_SAFETY as reasons. When the response is blocked, the candidate's finish_reason carries the reason: SAFETY, PROHIBITED_CONTENT, BLOCKLIST, SPII (sensitive personal information), RECITATION (too close to source material) or IMAGE_SAFETY.

Each carries safety_ratings: per category, a probability level and a blocked flag that names the category that tripped. Log those ratings on every request, not only on blocks. They are the raw material for calibration later.

Do not branch on what response.text does when content is missing; SDK convenience accessors have changed behaviour between versions. Classify the response explicitly:

RESPONSE_BLOCKS = {"SAFETY", "PROHIBITED_CONTENT", "BLOCKLIST", "SPII", "RECITATION", "IMAGE_SAFETY"}

def _name(enum_or_none):
    return getattr(enum_or_none, "name", None) or (str(enum_or_none) if enum_or_none else None)

def classify(resp):
    """Map a generate_content response to one outcome. Never read resp.text first."""
    fb = resp.prompt_feedback
    if fb is not None and fb.block_reason:
        return {"outcome": "prompt_blocked", "reason": _name(fb.block_reason),
                "ratings": _ratings(fb.safety_ratings)}
    if not resp.candidates:
        return {"outcome": "empty", "reason": None, "ratings": []}
    cand = resp.candidates[0]
    reason = _name(cand.finish_reason)
    if reason in RESPONSE_BLOCKS:
        return {"outcome": "response_blocked", "reason": reason,
                "ratings": _ratings(cand.safety_ratings)}
    text = "".join(p.text or "" for p in (cand.content.parts if cand.content else []))
    return {"outcome": "answered", "reason": reason, "text": text,
            "ratings": _ratings(cand.safety_ratings)}

def _ratings(rs):
    return [{"cat": _name(r.category), "p": _name(r.probability), "blocked": bool(r.blocked)}
            for r in (rs or [])]

Note that RECITATION is not a harm block. It means the output matched existing text too closely, and the right response is usually to retry with a prompt that asks for a summary or paraphrase, not to show a safety notice. Keeping reasons distinct lets each one get the right handling.

Streaming: the block that arrives after the text

With generate_content_stream, response filters run while tokens are produced. A stream can deliver several innocent chunks and then end with a chunk whose finish reason is SAFETY. By then your UI has rendered part of the answer. If you simply stop, the user sees a truncated paragraph that may already contain the problematic start, and you have no record that it happened.

Treat each chunk like a full response and run the same classifier on it. On a block, retract what was shown, replace it with a notice, and log how much text had already been displayed:

shown = []
for chunk in client.models.generate_content_stream(model=MODEL, contents=user_text,
                                                   config=types.GenerateContentConfig(safety_settings=SAFETY)):
    out = classify(chunk)
    if out["outcome"] in ("prompt_blocked", "response_blocked"):
        ui.retract(shown)                       # replace what was displayed, do not leave half an answer
        ui.show_notice(out["reason"])
        audit.log(request_id, out, partial_chars=sum(map(len, shown)))
        break
    if out.get("text"):
        shown.append(out["text"])
        ui.append(out["text"])

If retraction is unacceptable for your product, for example a voice interface where audio has already played, buffer a sentence or two before rendering. That costs some time-to-first-token but means most mid-stream blocks are caught before anything is shown. Measure how often blocks happen mid-stream before you choose; for many apps it is rare enough that retraction is fine.

Probability is not severity

A probability level answers "how likely is it that this text belongs to the category", not "how bad is it". Mild profanity can score a HIGH probability of harassment while a calmly worded piece of dangerous instruction scores only MEDIUM. Thresholds on probability therefore block a lot of harmless rudeness at strict settings and can miss serious content at loose ones.

Vertex AI's version of the API addresses this with a method field on each safety setting. Its HarmBlockMethod enum offers PROBABILITY and SEVERITY; the SDK reference describes SEVERITY as using both probability and severity scores and says the probability score is used when the method is not specified. On Vertex the ratings also carry numeric probability_score, severity and severity_score fields. Those scores are far better for calibration than four coarse levels, which is a real reason to prefer Vertex for products that need fine-grained thresholds. The Gemini Developer API returns the coarse levels only, so design your logging to accept both shapes.

Worked example: calibrating a pharmacology tutor

Suppose you are building a tutor for pharmacy students. Questions about lethal doses, drug interactions and overdose management are core curriculum. At BLOCK_MEDIUM_AND_ABOVE for dangerous content, a pilot shows 9% of legitimate questions blocked. Students learn to distrust the tool within a week.

Do not guess a new threshold. Build a labelled set of around 500 real questions: 450 legitimate curriculum queries and 50 that your policy says must not be answered, such as requests for help harming a specific person. Run the set with every category at BLOCK_NONE so content is returned regardless of score, and record the ratings per item; confirm in your logs that ratings are actually present for that setting before trusting the run. Then compute, offline, what each threshold would have done:

Dangerous-content thresholdLegitimate blockedDisallowed passed
BLOCK_LOW_AND_ABOVE71 of 450 (15.8%)0 of 50
BLOCK_MEDIUM_AND_ABOVE41 of 450 (9.1%)2 of 50
BLOCK_ONLY_HIGH6 of 450 (1.3%)9 of 50

These figures are illustrative, but the shape is typical: no single threshold gives both a low false-block rate and zero misses. The usable design is BLOCK_ONLY_HIGH for dangerous content plus your own check for the nine misses. Read them and you find a pattern the generic classifier cannot see: they all target a named or described individual. A small application-level classifier or rule for "harm to a specific person" catches that pattern without touching curriculum questions. The filter handles the generic tail; your layer handles your domain. Re-run the set whenever you change model, because ratings shift between model versions.

What the filter will never catch

The harm categories describe content. Several of the most damaging LLM failures are not content problems at all, and no threshold addresses them:

  • Prompt injection. Instructions hidden in a retrieved web page or document are polite, harmless-looking text. They score NEGLIGIBLE everywhere and can still redirect a tool-using agent. Defences live in your architecture; see jailbreak defence.
  • Leaking your own data. SPII covers some personal information, but not your customer's contract terms or your internal system prompt.
  • Unsafe actions. A model calling a refund tool with a wrong amount says nothing harmful. Validate tool arguments and treat output as untrusted input, as described in LLM output handling.
  • Domain policy. Medical, legal or financial advice rules, brand voice and regional law are yours to encode.
  • Deliberate evasion. Encodings, role-play and multi-turn build-ups are aimed precisely at classifiers like this one; the techniques are catalogued in moderation bypass.

Failure modes in production

  • Silent default. No safety_settings passed, so the adjustable layer is off on 2.5 and 3 models. Fix: set all four categories explicitly and assert it in a unit test on your config builder.
  • Crash on blocked prompt. Code indexes candidates[0] when there are none. Fix: classify first, always.
  • Refusals counted as answers. Dashboards show a 0.3% block rate while 6% of users get "I can't help with that". Fix: run a cheap refusal detector over answered text and report it as its own outcome.
  • Half-shown stream. A mid-stream block leaves partial text on screen. Fix: retraction or a short render buffer.
  • Retry storms. A client retries every blocked request, multiplying cost and producing the same block. Fix: retry only on RECITATION and transient errors, never on harm blocks.
  • Drift after model upgrade. Block rates move when the model id changes. Fix: re-run the labelled set as part of the upgrade checklist.

Operating it

Emit one structured event per request: request id, model id, the outcome from the classifier, the reason, and the per-category ratings. Do not log full prompts by default; log hashes, and keep full text only under a retention policy, since blocked prompts are often the most sensitive ones you receive. Chart outcomes per category and per surface, alert on sudden changes in either direction, and sample blocked and refused items into a review queue so humans read real cases each week. The broader review and calibration process is covered in content moderation architecture; the measurement side in LLM security evaluations.

Show users a specific, non-accusatory message for blocks, and give them a way to report a false block. Those reports are the cheapest source of new labelled data you will ever get.

What to do next

  1. Grep your code for every Gemini call and confirm each passes explicit safety_settings for all four categories.
  2. Add one classify() function and route every response, streamed or not, through it before reading text.
  3. Log outcome, reason and per-category ratings on every request, and add a refusal detector for answered text.
  4. Build a labelled set of a few hundred real prompts, run it at BLOCK_NONE, and pick thresholds from the measured trade-off.
  5. Write down which risks your own layer owns (injection, data leakage, tool actions, domain rules) and name the control for each.
  6. Add the labelled-set run to your model-upgrade checklist.
Key takeaway: Gemini's safety filters are three layers: fixed core protections, per-category thresholds you set, and the model's own refusals, which arrive as normal text. On Gemini 2.5 and 3 the adjustable thresholds default to Off, so set them explicitly. Classify every response from promptFeedback and finishReason before reading text, retract mid-stream blocks, calibrate thresholds on labelled data from your own traffic, and keep your own layer for injection, data leakage, tool actions and domain policy, which no content filter will ever cover.