Gemini scores every prompt and every candidate response against a set of harm categories and can block either one. Safety settings are the per-category thresholds that decide when a score becomes a block. In ADK Java they are one field on a config object, which makes them easy to set and easy to get wrong. The harder part is what happens after a block: the shape of the response changes, ADK translates it into its own types, and some of the information Gemini returned does not survive the translation.

This article covers configuring the thresholds per agent, exactly how a block reaches your code (read from the google-adk 1.11.0 bytecode), handling it in a callback so users get a useful answer, and calibrating thresholds with the raw genai client. For the safety model itself, the categories, the non-adjustable protections and trained refusals, start with Gemini safety settings, in depth.

Categories, thresholds and what they do not cover

A safety setting pairs a category with a threshold. In google-genai 1.75.0 the HarmCategory values include HARM_CATEGORY_HARASSMENT, HARM_CATEGORY_HATE_SPEECH, HARM_CATEGORY_SEXUALLY_EXPLICIT and HARM_CATEGORY_DANGEROUS_CONTENT, the four that both Gemini surfaces document as configurable. The enum also lists civic-integrity, jailbreak and image categories; Vertex AI documents its jailbreak classifier as a preview limited to particular models, so treat those as surface-specific and test before relying on them.

ThresholdEffect
BLOCK_LOW_AND_ABOVEBlock when the content is rated low risk or higher; the strictest setting
BLOCK_MEDIUM_AND_ABOVEBlock at medium or high
BLOCK_ONLY_HIGHBlock only at high
BLOCK_NONENever block on this category; Vertex AI still returns the scores
OFFTurn the filter off; Vertex AI documents that no safety metadata is returned

Three facts shape how you use them. First, defaults vary by surface and model generation: the Gemini API documentation says an unset threshold means Off for Gemini 2.5 and 3 models, while Vertex AI's page describes several defaults in different places. Do not inherit them; set every category explicitly so the behaviour is in your code and your review history. Second, HarmBlockMethod (SEVERITY or PROBABILITY) is a Vertex AI field, with severity documented as its default; leave it unset when you call the Gemini Developer API. Third, settings are not the whole system: protections such as child-safety blocks cannot be adjusted, and responses can stop with reasons like PROHIBITED_CONTENT, SPII or BLOCKLIST whatever your thresholds say.

Configuring profiles per agent

Settings travel in GenerateContentConfig, which LlmAgent.Builder accepts through generateContentConfig(...). Because each agent has its own config, a multi-agent app can be strict where users type freely and more permissive where the input is trusted, such as an internal agent that summarises incident reports full of violent language. Express that as named profiles rather than scattered builder calls:

public enum SafetyProfile {
  PUBLIC_CHAT(HarmBlockThreshold.Known.BLOCK_LOW_AND_ABOVE),
  STANDARD(HarmBlockThreshold.Known.BLOCK_MEDIUM_AND_ABOVE),
  TRUSTED_INTERNAL(HarmBlockThreshold.Known.BLOCK_ONLY_HIGH);

  private static final List<HarmCategory.Known> CATEGORIES = List.of(
      HarmCategory.Known.HARM_CATEGORY_HARASSMENT,
      HarmCategory.Known.HARM_CATEGORY_HATE_SPEECH,
      HarmCategory.Known.HARM_CATEGORY_SEXUALLY_EXPLICIT,
      HarmCategory.Known.HARM_CATEGORY_DANGEROUS_CONTENT);

  private final HarmBlockThreshold.Known threshold;
  SafetyProfile(HarmBlockThreshold.Known t) { this.threshold = t; }

  public List<SafetySetting> settings() {
    return CATEGORIES.stream()
        .map(cat -> SafetySetting.builder().category(cat).threshold(threshold).build())
        .toList();
  }
}

GenerateContentConfig cfg = GenerateContentConfig.builder()
    .temperature(0.3f)
    .safetySettings(SafetyProfile.PUBLIC_CHAT.settings())
    .build();

LlmAgent support = LlmAgent.builder()
    .name("support_chat")
    .model("gemini-2.5-flash")
    .instruction("Help customers with orders and returns.")
    .generateContentConfig(cfg)
    .afterModelCallback(SafetyBlockHandler::handle)
    .build();

A real profile would vary the threshold by category, for example stricter on harassment than on dangerous content for a chemistry-education app. Keep the mapping in one class, test it, and log the profile name with every request so an incident review can say which rules were in force. The other fields of the config are covered in Configuring Gemini models in ADK Java.

How a block reaches your code

Where a Gemini safety block surfaces in ADK Java 1.11.0LlmAgentgenerateContentConfigGeminiprompt + response filtersLlmResponse.Builderresponse(...)safetySettingsGenerateContentResponsePrompt blockederrorCode = blockReason, no finishReasonResponse blocked, no partsfinishReason = errorCode = SAFETYParts present or STOPcontent kept, no errorCodeafterModelCallback then Event.errorCode() at the runnersafetyRatings are not copied into LlmResponse, so per-category scores never reach callbacks.
ADK keeps content when parts are present or the candidate stopped normally; otherwise it fills errorCode from the finish reason or the prompt's block reason.

Gemini reports blocks in two places. A blocked prompt produces a response with no candidates and a promptFeedback.blockReason. A blocked response produces a candidate whose finishReason is SAFETY, whose content is withheld, and whose safetyRatings mark the category that tripped. ADK then converts the genai response into an LlmResponse in LlmResponse.Builder.response(...). Reading that method's bytecode in 1.11.0 gives four cases:

Gemini returnedfinishReason()errorCode()errorMessage()content
Candidate with parts, or finished STOPcandidate's reasonemptyemptykept
Candidate with no parts, not STOP (a response block)e.g. SAFETYsame as finishReasoncandidate's finishMessageempty
No candidates, promptFeedback present (a prompt block)emptyblockReason as a stringblockReasonMessageempty
No candidates, no feedbackempty"Unknown error.""Unknown error."empty

Three consequences follow. The per-category safetyRatings are not copied: LlmResponse has no field for them, so no ADK callback can see which category fired or how high it scored. A prompt block and a response block can both show an error code whose text is SAFETY; tell them apart by finishReason(), which is present only for the response block. And because the prompt-block code is built with new FinishReason(blockReason.toString()), values that are not in FinishReason.Known, such as JAILBREAK or MODEL_ARMOR, come back with knownEnum() equal to FINISH_REASON_UNSPECIFIED. Classify on toString(), never on the known enum alone.

Handling blocks in an afterModelCallback

An afterModelCallback runs on every model response before ADK turns it into an event. Returning a replacement LlmResponse swaps what the user sees; returning Maybe.empty() keeps the original. That is the place to turn a bare block into a useful answer and to count blocks:

public final class SafetyBlockHandler {
  enum Kind { PROMPT_BLOCK, RESPONSE_BLOCK, OTHER_STOP, NONE }

  static Kind classify(LlmResponse r) {
    if (r.errorCode().isEmpty()) return Kind.NONE;
    String code = r.errorCode().get().toString();
    if (r.finishReason().isEmpty() && !code.equals("Unknown error.")) return Kind.PROMPT_BLOCK;
    if (Set.of("SAFETY", "PROHIBITED_CONTENT", "SPII", "BLOCKLIST").contains(code)) {
      return Kind.RESPONSE_BLOCK;
    }
    return Kind.OTHER_STOP;                       // MAX_TOKENS with no parts, MALFORMED_FUNCTION_CALL...
  }

  public static Maybe<LlmResponse> handle(CallbackContext ctx, LlmResponse r) {
    Kind kind = classify(r);
    if (kind == Kind.NONE || kind == Kind.OTHER_STOP) return Maybe.empty();
    String code = r.errorCode().get().toString();
    Metrics.counter("adk.safety.blocks", "kind", kind.name(), "code", code).increment();
    log.warn("safety block kind={} code={} agent={} invocation={}",
        kind, code, ctx.agentName(), ctx.invocationId());
    String reply = kind == Kind.PROMPT_BLOCK
        ? "I can't help with that request as written. Could you rephrase it?"
        : "I started an answer but couldn't complete it. Could you ask in a different way?";
    return Maybe.just(LlmResponse.builder()
        .content(Content.builder().role("model").parts(Part.fromText(reply)).build())
        .customMetadata(List.of(CustomMetadata.builder()
            .key("safety_block").stringValue(kind + ":" + code).build()))
        .build());
  }
}

The replacement drops the error code, so downstream code sees a normal answer, and the custom metadata keeps the fact of the block on the event for analytics. Logging the code string, the agent and the invocation ID is enough to join with request logs later; do not log the blocked prompt itself into a general-purpose log. If you would rather surface the block to a client UI than paraphrase it, skip the callback and read Event.errorCode() on the events runAsync emits. For the full callback order and the other hooks, see ADK Java callbacks.

Testing the conversion and the handler

The conversion rules above are easy to verify, because LlmResponse.create accepts a genai response you build by hand. Pin them in unit tests so an ADK upgrade that changes the mapping fails your build instead of your users:

@Test void promptBlockKeepsUnknownReasonAsString() {
  GenerateContentResponse raw = GenerateContentResponse.builder()
      .promptFeedback(GenerateContentResponsePromptFeedback.builder()
          .blockReason(BlockedReason.Known.JAILBREAK))
      .build();
  LlmResponse r = LlmResponse.create(raw);

  assertEquals("JAILBREAK", r.errorCode().orElseThrow().toString());
  assertEquals(FinishReason.Known.FINISH_REASON_UNSPECIFIED, r.errorCode().get().knownEnum());
  assertTrue(r.finishReason().isEmpty());
  assertEquals(Kind.PROMPT_BLOCK, SafetyBlockHandler.classify(r));
}

@Test void emptySafetyCandidateBecomesResponseBlock() {
  GenerateContentResponse raw = GenerateContentResponse.builder()
      .candidates(Candidate.builder().finishReason(FinishReason.Known.SAFETY))
      .build();
  LlmResponse r = LlmResponse.create(raw);

  assertEquals("SAFETY", r.errorCode().orElseThrow().toString());
  assertEquals(Kind.RESPONSE_BLOCK, SafetyBlockHandler.classify(r));
}

Add a third test for a candidate that has text and finishes with SAFETY: it should keep its content and carry no error code, which is the streaming case below. These tests run in milliseconds with no network, and together with an end-to-end check against the real API on a known-blocked prompt they cover both your code and the service.

Treat safety profiles as configuration under change control. A threshold change alters what users can get from the product, so it deserves a reviewed commit, a note of the calibration run that justified it and a staged rollout, not a hot edit in a console.

Streaming: blocks after partial text

With streaming, partial chunks arrive as separate responses. A chunk with text keeps its content even if a later chunk is blocked, because the content-kept branch only needs parts to be present. So the user may already have read half a sentence when the final chunk arrives with no parts and a SAFETY error code. The client must treat that terminal error as a retraction: clear or collapse the partial text, then show the replacement message. A server that relays tokens straight to a browser should hold a small buffer or mark streamed text as provisional until the turn completes. The mechanics of ADK streaming are in Gemini streaming in ADK Java.

Calibrating thresholds with the genai client

Choosing thresholds needs the scores, and ADK drops them. Run calibration outside the agent with the genai client, against a labelled set of prompts drawn from real traffic: requests you must serve, such as a nurse asking about overdose thresholds, and requests you must refuse. On Vertex AI, BLOCK_NONE returns the ratings without blocking, which is exactly what calibration wants.

Client client = Client.builder().vertexAI(true).project(PROJECT).location(LOCATION).build();
GenerateContentConfig probe = GenerateContentConfig.builder()
    .safetySettings(allCategoriesAt(HarmBlockThreshold.Known.BLOCK_NONE))
    .maxOutputTokens(256)
    .build();

for (LabelledPrompt lp : evalSet) {
  GenerateContentResponse res = client.models.generateContent(MODEL, lp.text(), probe);
  List<SafetyRating> ratings = res.candidates().flatMap(cs -> cs.stream().findFirst())
      .flatMap(Candidate::safetyRatings).orElse(List.of());
  for (SafetyRating r : ratings) {
    csv.write(lp.id(), lp.mustServe(), r.category(), r.probability(), r.severity(),
              r.probabilityScore().orElse(null), r.severityScore().orElse(null));
  }
}

allCategoriesAt is a small helper that builds the four settings at one threshold. From the CSV, compute for each candidate threshold how many must-serve prompts would be blocked (false refusals) and how many must-refuse prompts would pass. Response ratings matter as much as prompt ratings, since a harmless question can draw a flagged answer, so rate the candidates too.

Worked example: a pharmacy support agent

A pharmacy retailer's support agent runs with PUBLIC_CHAT. Suppose the team's dashboards show blocks on about 2 percent of turns, and the logged codes show they are almost all response blocks on dangerous content, mostly dosage questions. (These figures are illustrative; your own logs supply the real ones.) Suppose a calibration run over 400 labelled prompts shows that relaxing dangerous content to BLOCK_MEDIUM_AND_ABOVE clears most of the false refusals while the must-refuse set stays blocked by the model's own refusals and the stricter categories. The team ships a new profile that changes that one category, keeps the other three at the strictest level, and watches the block counter: the rate falls, and the remaining blocks are reviewed weekly.

The lesson is the order of work: measure what is blocked and why, test a change offline against labelled data, change one category at a time, and keep the counter running.

Failure modes

  • Relying on defaults. A model upgrade or a move between the Gemini API and Vertex AI can change unset thresholds silently. Set all four.
  • Switching on the known enum. Blocked-reason values outside FinishReason.Known read as unspecified; classify on the string.
  • Treating an empty answer as success. Without a handler a blocked turn can reach the client as an event with no text. Check errorCode() everywhere you render events.
  • Using OFF for calibration. On Vertex AI OFF returns no scores; use BLOCK_NONE to measure.
  • Expecting settings to stop trained refusals. A refusal is ordinary text with finish reason STOP; no threshold changes it, and it will not show up in block counts.
  • Treating settings as the guardrail. They know nothing about your domain; keep the application checks described in ADK Java guardrails.

Trade-offs

Strict thresholds protect users and the brand at the cost of false refusals, which frustrate exactly the users with legitimate but sensitive questions; permissive thresholds reverse the balance. Paraphrasing a block in a callback gives users a smooth experience but hides the event from the client; surfacing the error code lets the UI explain itself but leaks implementation detail. Per-agent profiles fit each agent's risk, but every extra profile is configuration to review and test, so keep the set small and named. Finally, running BLOCK_NONE and applying your own classifier on the returned scores gives full control and full responsibility, while Gemini's built-in filters are less flexible but maintained for you. Most teams should start with explicit built-in thresholds and add their own layer only for domain rules the filters cannot express.

What to do next

  1. Write down the safety profile each agent in your app should run, and why.
  2. Encode the profiles in one class and set all four categories explicitly in every agent's GenerateContentConfig.
  3. Add an afterModelCallback that classifies blocks by string, replaces the reply and increments a counter.
  4. Make your streaming client handle a terminal error code as a retraction of partial text.
  5. Build a labelled prompt set from real traffic and run the calibration harness with BLOCK_NONE on Vertex AI.
  6. Review block counts by code and agent weekly, and change one category at a time.
Key takeaway: Safety settings are per-category thresholds carried in each agent's GenerateContentConfig; set all of them explicitly, because defaults differ by surface and model. In ADK Java 1.11.0 a prompt block arrives as an errorCode with no finishReason, a response block as matching finishReason and errorCode, and the per-category ratings are dropped. Classify blocks by their string code in an afterModelCallback, retract streamed text on a terminal block, and calibrate thresholds offline with the genai client and BLOCK_NONE.