If you build on the Gemini API, safety filtering is not one thing you switch on. Three separate mechanisms can stop or change an answer, they report themselves in different places in the response, and one of them does not report itself at all. Teams that treat the filter as a single boolean ship apps that crash on an empty response.text, show users half an answer that was cut mid-stream, or quietly loosen a threshold until something embarrassing gets through.
This article explains the filter from the integrator's side: what each layer does, which settings you control, how to read every block signal, how to pick thresholds from your own data rather than by feel, and which risks you still own after the filter has done its job. Specifics were checked against Google's safety-settings guide and API reference on 2026-10-01; the API evolves, so verify enum names against the reference before you depend on them.
Three layers, not one switch
The first layer is a set of built-in protections against core harms. Google's guide gives content that endangers child safety as the example and says these are always blocked and cannot be adjusted. You cannot configure them, and your app should treat them as a hard edge.
The second layer is the adjustable classifier. Every prompt and every candidate response is scored per harm category, and you choose, per category, how high the score must be before the request or response is blocked. This is what people usually mean by "safety settings".
The third layer is the model itself. Gemini is trained to decline some requests. A trained refusal is not a block: the response arrives normally, with text such as "I can't help with that" and a finish reason of STOP. No setting turns this off, and none of the block fields mention it. If you count only flagged blocks, you undercount how often users fail to get an answer.
Outside all three sits your application layer: your own input checks, output handling and domain policy. The filter knows nothing about your product, your users or your data, so that layer is never optional.
The knobs you actually have
The guide lists four adjustable categories: HARM_CATEGORY_HARASSMENT, HARM_CATEGORY_HATE_SPEECH, HARM_CATEGORY_SEXUALLY_EXPLICIT and HARM_CATEGORY_DANGEROUS_CONTENT. The API reference also names further categories, such as civic integrity and jailbreak; whether a given model honours them varies, so test before you rely on one.
Each category takes a HarmBlockThreshold. The classifier reports a probability level of NEGLIGIBLE, LOW, MEDIUM or HIGH, and the threshold says which levels are blocked:
| Threshold | Blocks when probability is | Typical use |
|---|---|---|
BLOCK_LOW_AND_ABOVE | LOW, MEDIUM or HIGH | Children's products, strict brand surfaces |
BLOCK_MEDIUM_AND_ABOVE | MEDIUM or HIGH | General consumer chat |
BLOCK_ONLY_HIGH | HIGH | Professional tools where false blocks are costly |
BLOCK_NONE | never; content shown regardless of probability | Measurement runs, vetted internal tools |
OFF | filter turned off for the category | Same as above, where you own the downstream check |
HARM_BLOCK_THRESHOLD_UNSPECIFIED | uses the default | Avoid; be explicit |
The most important fact in the table is the default. The guide (checked 2026-10-01) says that when no threshold is set, the default block threshold is Off for Gemini 2.5 and 3 models. If you never pass safety_settings, the adjustable layer is not blocking anything on those models, and you are relying on the core protections, the model's own refusals and whatever you built. Many teams assume the opposite. Set every category explicitly so the behaviour is visible in code review and does not shift when you change model.
from google import genai
from google.genai import types
MODEL = "your-gemini-model-id" # pick a current id from the models page; do not hard-code a stale one
client = genai.Client() # reads GEMINI_API_KEY from the environment
SAFETY = [
types.SafetySetting(category=types.HarmCategory.HARM_CATEGORY_HARASSMENT,
threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE),
types.SafetySetting(category=types.HarmCategory.HARM_CATEGORY_HATE_SPEECH,
threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE),
types.SafetySetting(category=types.HarmCategory.HARM_CATEGORY_SEXUALLY_EXPLICIT,
threshold=types.HarmBlockThreshold.BLOCK_LOW_AND_ABOVE),
types.SafetySetting(category=types.HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT,
threshold=types.HarmBlockThreshold.BLOCK_ONLY_HIGH),
]
resp = client.models.generate_content(
model=MODEL,
contents=user_text,
config=types.GenerateContentConfig(safety_settings=SAFETY),
)
Where the verdict shows up
A blocked prompt and a blocked response look different. When the prompt is blocked, prompt_feedback.block_reason is set and there are no candidates to read. The reference lists SAFETY, OTHER, BLOCKLIST, PROHIBITED_CONTENT and IMAGE_SAFETY as reasons. When the response is blocked, the candidate's finish_reason carries the reason: SAFETY, PROHIBITED_CONTENT, BLOCKLIST, SPII (sensitive personal information), RECITATION (too close to source material) or IMAGE_SAFETY.
Each carries safety_ratings: per category, a probability level and a blocked flag that names the category that tripped. Log those ratings on every request, not only on blocks. They are the raw material for calibration later.
Do not branch on what response.text does when content is missing; SDK convenience accessors have changed behaviour between versions. Classify the response explicitly:
RESPONSE_BLOCKS = {"SAFETY", "PROHIBITED_CONTENT", "BLOCKLIST", "SPII", "RECITATION", "IMAGE_SAFETY"}
def _name(enum_or_none):
return getattr(enum_or_none, "name", None) or (str(enum_or_none) if enum_or_none else None)
def classify(resp):
"""Map a generate_content response to one outcome. Never read resp.text first."""
fb = resp.prompt_feedback
if fb is not None and fb.block_reason:
return {"outcome": "prompt_blocked", "reason": _name(fb.block_reason),
"ratings": _ratings(fb.safety_ratings)}
if not resp.candidates:
return {"outcome": "empty", "reason": None, "ratings": []}
cand = resp.candidates[0]
reason = _name(cand.finish_reason)
if reason in RESPONSE_BLOCKS:
return {"outcome": "response_blocked", "reason": reason,
"ratings": _ratings(cand.safety_ratings)}
text = "".join(p.text or "" for p in (cand.content.parts if cand.content else []))
return {"outcome": "answered", "reason": reason, "text": text,
"ratings": _ratings(cand.safety_ratings)}
def _ratings(rs):
return [{"cat": _name(r.category), "p": _name(r.probability), "blocked": bool(r.blocked)}
for r in (rs or [])]Note that RECITATION is not a harm block. It means the output matched existing text too closely, and the right response is usually to retry with a prompt that asks for a summary or paraphrase, not to show a safety notice. Keeping reasons distinct lets each one get the right handling.
Streaming: the block that arrives after the text
With generate_content_stream, response filters run while tokens are produced. A stream can deliver several innocent chunks and then end with a chunk whose finish reason is SAFETY. By then your UI has rendered part of the answer. If you simply stop, the user sees a truncated paragraph that may already contain the problematic start, and you have no record that it happened.
Treat each chunk like a full response and run the same classifier on it. On a block, retract what was shown, replace it with a notice, and log how much text had already been displayed:
shown = []
for chunk in client.models.generate_content_stream(model=MODEL, contents=user_text,
config=types.GenerateContentConfig(safety_settings=SAFETY)):
out = classify(chunk)
if out["outcome"] in ("prompt_blocked", "response_blocked"):
ui.retract(shown) # replace what was displayed, do not leave half an answer
ui.show_notice(out["reason"])
audit.log(request_id, out, partial_chars=sum(map(len, shown)))
break
if out.get("text"):
shown.append(out["text"])
ui.append(out["text"])If retraction is unacceptable for your product, for example a voice interface where audio has already played, buffer a sentence or two before rendering. That costs some time-to-first-token but means most mid-stream blocks are caught before anything is shown. Measure how often blocks happen mid-stream before you choose; for many apps it is rare enough that retraction is fine.
Probability is not severity
A probability level answers "how likely is it that this text belongs to the category", not "how bad is it". Mild profanity can score a HIGH probability of harassment while a calmly worded piece of dangerous instruction scores only MEDIUM. Thresholds on probability therefore block a lot of harmless rudeness at strict settings and can miss serious content at loose ones.
Vertex AI's version of the API addresses this with a method field on each safety setting. Its HarmBlockMethod enum offers PROBABILITY and SEVERITY; the SDK reference describes SEVERITY as using both probability and severity scores and says the probability score is used when the method is not specified. On Vertex the ratings also carry numeric probability_score, severity and severity_score fields. Those scores are far better for calibration than four coarse levels, which is a real reason to prefer Vertex for products that need fine-grained thresholds. The Gemini Developer API returns the coarse levels only, so design your logging to accept both shapes.
Worked example: calibrating a pharmacology tutor
Suppose you are building a tutor for pharmacy students. Questions about lethal doses, drug interactions and overdose management are core curriculum. At BLOCK_MEDIUM_AND_ABOVE for dangerous content, a pilot shows 9% of legitimate questions blocked. Students learn to distrust the tool within a week.
Do not guess a new threshold. Build a labelled set of around 500 real questions: 450 legitimate curriculum queries and 50 that your policy says must not be answered, such as requests for help harming a specific person. Run the set with every category at BLOCK_NONE so content is returned regardless of score, and record the ratings per item; confirm in your logs that ratings are actually present for that setting before trusting the run. Then compute, offline, what each threshold would have done:
| Dangerous-content threshold | Legitimate blocked | Disallowed passed |
|---|---|---|
| BLOCK_LOW_AND_ABOVE | 71 of 450 (15.8%) | 0 of 50 |
| BLOCK_MEDIUM_AND_ABOVE | 41 of 450 (9.1%) | 2 of 50 |
| BLOCK_ONLY_HIGH | 6 of 450 (1.3%) | 9 of 50 |
These figures are illustrative, but the shape is typical: no single threshold gives both a low false-block rate and zero misses. The usable design is BLOCK_ONLY_HIGH for dangerous content plus your own check for the nine misses. Read them and you find a pattern the generic classifier cannot see: they all target a named or described individual. A small application-level classifier or rule for "harm to a specific person" catches that pattern without touching curriculum questions. The filter handles the generic tail; your layer handles your domain. Re-run the set whenever you change model, because ratings shift between model versions.
What the filter will never catch
The harm categories describe content. Several of the most damaging LLM failures are not content problems at all, and no threshold addresses them:
- Prompt injection. Instructions hidden in a retrieved web page or document are polite, harmless-looking text. They score NEGLIGIBLE everywhere and can still redirect a tool-using agent. Defences live in your architecture; see jailbreak defence.
- Leaking your own data.
SPIIcovers some personal information, but not your customer's contract terms or your internal system prompt. - Unsafe actions. A model calling a refund tool with a wrong amount says nothing harmful. Validate tool arguments and treat output as untrusted input, as described in LLM output handling.
- Domain policy. Medical, legal or financial advice rules, brand voice and regional law are yours to encode.
- Deliberate evasion. Encodings, role-play and multi-turn build-ups are aimed precisely at classifiers like this one; the techniques are catalogued in moderation bypass.
Failure modes in production
- Silent default. No
safety_settingspassed, so the adjustable layer is off on 2.5 and 3 models. Fix: set all four categories explicitly and assert it in a unit test on your config builder. - Crash on blocked prompt. Code indexes
candidates[0]when there are none. Fix: classify first, always. - Refusals counted as answers. Dashboards show a 0.3% block rate while 6% of users get "I can't help with that". Fix: run a cheap refusal detector over answered text and report it as its own outcome.
- Half-shown stream. A mid-stream block leaves partial text on screen. Fix: retraction or a short render buffer.
- Retry storms. A client retries every blocked request, multiplying cost and producing the same block. Fix: retry only on
RECITATIONand transient errors, never on harm blocks. - Drift after model upgrade. Block rates move when the model id changes. Fix: re-run the labelled set as part of the upgrade checklist.
Operating it
Emit one structured event per request: request id, model id, the outcome from the classifier, the reason, and the per-category ratings. Do not log full prompts by default; log hashes, and keep full text only under a retention policy, since blocked prompts are often the most sensitive ones you receive. Chart outcomes per category and per surface, alert on sudden changes in either direction, and sample blocked and refused items into a review queue so humans read real cases each week. The broader review and calibration process is covered in content moderation architecture; the measurement side in LLM security evaluations.
Show users a specific, non-accusatory message for blocks, and give them a way to report a false block. Those reports are the cheapest source of new labelled data you will ever get.
What to do next
- Grep your code for every Gemini call and confirm each passes explicit
safety_settingsfor all four categories. - Add one
classify()function and route every response, streamed or not, through it before reading text. - Log outcome, reason and per-category ratings on every request, and add a refusal detector for answered text.
- Build a labelled set of a few hundred real prompts, run it at
BLOCK_NONE, and pick thresholds from the measured trade-off. - Write down which risks your own layer owns (injection, data leakage, tool actions, domain rules) and name the control for each.
- Add the labelled-set run to your model-upgrade checklist.