PaLM is gone. The Gemini Developer API decommissioned the PaLM API on 15 August 2024, and Vertex AI followed on its own, later schedule, cutting off new projects first and removing access for all projects in 2025. If a Java service still contains a text-bison or chat-bison call, that code path is already failing, or it sits behind a feature flag nobody has turned on for a year.

One point is worth stating up front, because it trips people up: the Agent Development Kit for Java never had a PaLM model class. ADK Java arrived after PaLM was retired and speaks Gemini natively. So "migrating from PaLM to Gemini in ADK Java" means taking a Java service that called PaLM through the Vertex AI prediction API (or a framework wrapper over it) and rebuilding its model access on ADK Java with Gemini. This article covers the concept mapping, a worked rewrite behind a port, an evaluation gate that decides when the new path is good enough, and how to make the next migration cheaper. That next migration is not hypothetical: Google's lifecycle pages already list the Gemini 2.5 models for retirement in October 2026.

Advertisement

What actually changes

PaLM and Gemini differ in more than the model name. The request shape, the conversation model, the parameters, the safety system and the response structure all changed. The table maps each PaLM-era concept to where it lives in Gemini and in ADK Java.

PaLM conceptGemini APIIn ADK Java
text-bison promptA user Content with text partsThe user message passed to runAsync
chat-bison contextSystem instructionLlmAgent.builder().instruction(...)
chat-bison examples (input/output pairs)Few-shot text in the instruction, or prior turnsInstruction text, or seeded session history
messages with authorContents with role user or modelSession events, managed by the runner
temperature, topP, topK, maxOutputTokensSame names in GenerateContentConfig.generateContentConfig(...) on the agent
candidateCount, stopSequencesSame names in GenerateContentConfigSame; request one candidate per call
safetyAttributes scores and blockedsafetySettings thresholds in, safety ratings and finish reason outSafety settings in the content config
predictions[0].contentCandidates, each with a list of partsEvent objects; finalResponse() marks the answer
PaLM-tuned modelsNot portableRe-tune on a Gemini base or replace with prompting
textembedding-gecko vectorsNew embedding model, new vector spaceRe-embed the whole corpus

Two rows deserve emphasis. Tuned PaLM models do not carry over: there is no conversion, and the training data is the only reusable asset. And embeddings from different models are not comparable. If your retrieval index holds vectors from a PaLM-era embedding model, queries embedded with a new model will return nonsense against it. Re-embedding the corpus is a data migration with its own backfill, dual-index and cutover plan; schedule it separately from the generation change.

The starting point: a typical PaLM-era Java call

Most Java services called PaLM through the Vertex AI Java client's prediction API, with the request body as protobuf Value objects built from JSON. A representative (historical) example:

// Historical: PaLM text-bison through the Vertex AI prediction API. These models no longer serve.
EndpointName endpoint = EndpointName.ofProjectLocationPublisherModelName(
    project, "us-central1", "google", "text-bison@002");

Value instance = toValue(Map.of("prompt", promptTemplate.formatted(customerQuestion)));
Value parameters = toValue(Map.of(
    "temperature", 0.2, "maxOutputTokens", 256, "topK", 40, "topP", 0.95));

PredictResponse response = predictionClient.predict(endpoint, List.of(instance), parameters);
Struct prediction = response.getPredictions(0).getStructValue();
String answer = prediction.getFieldsOrThrow("content").getStringValue();

Note what is wrong with this from a migration point of view, not only that the model is gone. The model name is a literal in business code. The prompt template mixes instructions and user input in one string. The response is parsed by field name, with no check for a blocked or empty result. Typically this pattern was copied into several classes. The rewrite fixes the structure first, so that this is the last migration that requires touching business code.

Advertisement

The architecture: put a port in front of the model

Define a small interface that expresses what the business code needs, in your own types, and make the ADK agent one implementation of it. The business code stops knowing which model, SDK or prompt format is in use.

From a hand-rolled PaLM client to an ADK Java agent behind a portBefore (PaLM era)After (ADK Java + Gemini)Service codebuilds prompt strings by handPredictionServiceClientinstances: prompt / context / examplestext-bison / chat-bisonendpoints removed; calls failParse predictions[0].contentsafetyAttributes.blockedService codeunchanged callersTextModel portyour interface, your recordsLlmAgent + Runnerinstruction, config, sessionsGemini modelID from configurationGolden-set replayquality, refusals, tokens, latencyModel name, prompt format and parsingwere spread through business code.
Before, model details leaked into callers. After, callers depend on a port, the ADK agent sits behind it with its model ID in configuration, and a replay harness runs the same golden set against any implementation.
public interface TextModel {
  String startConversation(String userId);
  Completion complete(String userId, String conversationId, String userText);
}

public record Completion(String text, int promptTokens, int outputTokens, boolean empty) {}

The port is deliberately smaller than either SDK. It has no temperature, no safety enum and no model name, because callers should not choose those. They belong to the implementation's configuration, where an operator can change them without a code review of every caller.

The rewrite: an ADK Java implementation

Add the ADK dependency (com.google.adk:google-adk; 1.11.0 is the version the ADK Java quickstart showed when this was written) and build an LlmAgent. The PaLM chat context becomes the agent's instruction, the parameters move into a GenerateContentConfig, and the model ID comes from configuration.

import com.google.adk.agents.LlmAgent;
import com.google.adk.agents.RunConfig;
import com.google.adk.runner.InMemoryRunner;
import com.google.adk.sessions.Session;
import com.google.genai.types.Content;
import com.google.genai.types.GenerateContentConfig;
import com.google.genai.types.Part;

public final class AdkTextModel implements TextModel {
  private final InMemoryRunner runner;              // swap in a persistent session service for production
  private final RunConfig runConfig = RunConfig.builder().build();

  public AdkTextModel(ModelSettings s) {
    LlmAgent agent = LlmAgent.builder()
        .name("order_support")
        .description("Answers customer questions about orders.")
        .model(s.modelId())                         // from configuration, never a literal here
        .instruction(s.instruction())               // was the chat-bison context, plus few-shot examples
        .generateContentConfig(GenerateContentConfig.builder()
            .temperature(s.temperature())
            .topP(s.topP())
            .maxOutputTokens(s.maxOutputTokens())
            .stopSequences(s.stopSequences())
            .safetySettings(s.safetySettings())
            .build())
        .build();
    this.runner = new InMemoryRunner(agent);
  }

  @Override
  public String startConversation(String userId) {
    Session session = runner.sessionService().createSession(runner.appName(), userId).blockingGet();
    return session.id();
  }

  @Override
  public Completion complete(String userId, String conversationId, String userText) {
    Content message = Content.fromParts(Part.fromText(userText));
    StringBuilder text = new StringBuilder();
    int[] tokens = new int[2];
    runner.runAsync(userId, conversationId, message, runConfig).blockingForEach(event -> {
      if (event.finalResponse()) {
        text.append(event.stringifyContent());
      }
      event.usageMetadata().ifPresent(u -> {
        tokens[0] += u.promptTokenCount().orElse(0);
        tokens[1] += u.candidatesTokenCount().orElse(0);
      });
    });
    return new Completion(text.toString(), tokens[0], tokens[1], text.length() == 0);
  }
}

Conversation history changed the most. With chat-bison the caller resent the whole messages list on every request. With ADK, the runner records each turn as an event in a session and builds the model's history from it, so the caller sends only the new user message plus a session id. That removes a class of bugs (truncated or reordered history) and adds a new responsibility: an in-memory session service loses every conversation on restart, so production needs a persistent one.

Few-shot examples from chat-bison's examples field have no dedicated slot. The simplest faithful mapping is to render them into the instruction under a heading such as "Examples", in the same input and output format the model should follow. Keep them short; on Gemini, clear instructions and an output schema often replace examples that PaLM needed.

Parameters do not transfer one to one

Even where names match, values do not. A temperature of 0.2 on text-bison and 0.2 on a Gemini model are different behaviours, because sampling acts on different models' distributions. Treat every PaLM parameter value as a starting guess and re-tune it against the evaluation set below.

Three specific traps. First, output budgets: on Gemini models that use internal reasoning, thinking tokens can count against the output budget, so a maxOutputTokens value carried over from PaLM can truncate answers or produce empty ones with a token-limit finish reason. Check the model's documentation, and either raise the limit or set a thinking budget; the Gemini 2.5 features article covers thinking configuration in ADK Java. Second, topK support and defaults vary by model, so set it only if evaluation shows it helps. Third, candidateCount above 1 costs tokens for every candidate while an ADK agent reads one; if you used multiple PaLM candidates for reranking, implement that explicitly with separate calls.

Token counts also differ between tokenizers, so per-request cost and context limits must be re-measured, not converted. See token counting across models.

Safety and response handling

PaLM returned safety scores and a blocked flag alongside the content. Gemini takes per-category thresholds in the request and, when a response is blocked, ends the candidate with a safety finish reason and little or no text. Code that only checked whether content was non-empty will now see empty answers and pass them to users. Configure thresholds explicitly, using the forms the Google Gen AI Java SDK documents:

List<SafetySetting> safety = List.of(
    SafetySetting.builder()
        .category(HarmCategory.Known.HARM_CATEGORY_HATE_SPEECH)
        .threshold(HarmBlockThreshold.Known.BLOCK_ONLY_HIGH)
        .build(),
    SafetySetting.builder()
        .category(HarmCategory.Known.HARM_CATEGORY_DANGEROUS_CONTENT)
        .threshold(HarmBlockThreshold.Known.BLOCK_LOW_AND_ABOVE)
        .build());

Then make the empty result a first-class outcome. The Completion.empty flag in the port lets callers show a fallback message, and a metric on it tells you when a prompt change or a new model version starts blocking legitimate traffic. The categories differ from PaLM's, so a threshold policy cannot be copied across; decide it again with the people who own content policy. Broader guardrail design is covered in safety in ADK Java.

Parsing changes too. A Gemini candidate holds a list of parts, which can include function calls and thought summaries as well as text. If you previously extracted JSON from free text with a regular expression, replace that with a response schema or the agent's output schema, which gives the model a contract instead of a hope.

The evaluation gate

Never cut over on the strength of a few manual prompts. Build a golden set from production traffic: a few hundred real inputs, de-identified, covering common cases, known hard cases and inputs that should be refused. For each, record what a good answer must contain rather than one exact expected string, because no two models phrase things identically.

record Case(String input, List<String> mustMention, boolean shouldRefuse) {}
record Score(int passed, int failed, int emptyOrBlocked, long p95Millis, long outputTokens) {}

Score replay(TextModel model, List<Case> cases) {
  int passed = 0, failed = 0, empty = 0; long outTokens = 0;
  List<Long> latencies = new ArrayList<>();
  for (Case c : cases) {
    String conv = model.startConversation("eval");
    long t0 = System.nanoTime();
    Completion out = model.complete("eval", conv, c.input());
    latencies.add((System.nanoTime() - t0) / 1_000_000);
    outTokens += out.outputTokens();
    if (out.empty()) empty++;
    boolean ok = c.shouldRefuse() ? looksLikeRefusal(out.text())
                                  : c.mustMention().stream().allMatch(m -> containsIgnoreCase(out.text(), m));
    if (ok) passed++; else failed++;
  }
  return new Score(passed, failed, empty, percentile(latencies, 95), outTokens);
}

Set pass criteria before you look at results: for example, pass rate at least the old system's historical rate, unexpected empty or blocked answers under 1%, p95 latency within budget, and cost per thousand requests within budget. Keyword checks are crude; add an LLM-as-judge rubric for the answers that pass them, and sample a slice for human review. Run the gate in CI on every prompt or model change, not only during the migration.

Rolling out

Because the PaLM endpoints no longer answer, there is no old system to shadow against in most cases. Use recorded outputs instead: the golden set holds what PaLM used to answer, and that history is your baseline. Roll out behind a flag by percentage of traffic, watch the empty-answer rate, user feedback and latency, and keep the flag until a full week of traffic has passed cleanly.

Record the model version that served each response. Aliases such as gemini-flash-latest are convenient and move underneath you; the ADK documentation also warns that this selector may not work against regional endpoints. Pin an explicit model version in production, and use the alias only in a canary environment where an upgrade is something you want to see early.

Failure modes

FailureCauseFix
Empty answers in productionSafety block or output budget spent on thinkingCheck finish reasons, raise the limit or set a thinking budget
Retrieval returns irrelevant passagesOld embeddings queried with a new modelRe-embed the corpus; never mix vector spaces
Conversations forget context after deploysIn-memory session servicePersistent session storage
JSON parse errorsRegex over free textResponse or output schema
Model works in one region, fails in anotherAlias or model not available thereExplicit versions, checked per region
Outage on the next retirementModel ID hard-coded againID in configuration, retirement dates tracked

What to do next

  1. Search the codebase for text-bison, chat-bison, code-bison, textembedding-gecko and PredictionServiceClient, and list every call site.
  2. Define a TextModel port in your own types and move every call site onto it before changing models.
  3. Implement the port with an ADK Java LlmAgent: context into the instruction, parameters into GenerateContentConfig, model ID from configuration.
  4. Build a golden set from recorded traffic, agree pass criteria, and run the replay gate in CI.
  5. Plan the embedding re-index as a separate migration with a dual index and cutover.
  6. Move sessions to persistent storage and add metrics for empty answers, tokens and latency per model version; a BaseLlm decorator is one place to collect them.
  7. Put your current Gemini model's published retirement date on the team calendar and rehearse the next swap with the same gate.
Key takeaway: ADK Java never ran PaLM, so this migration is a rebuild of model access, not a model rename. Put a small port in front of the model, implement it with an LlmAgent whose instruction holds the old context and examples and whose model ID lives in configuration, re-tune parameters rather than copying them, handle blocked or empty answers explicitly, re-embed any PaLM-era vectors, and cut over only when a golden-set replay gate passes. Done this way, the next retirement, already scheduled for Gemini 2.5, becomes a configuration change and a gate run.