In Google's Agent Development Kit for Java, the prompt is not one string. An LlmAgent carries an instruction, optionally a global instruction, a description that parent agents use for routing, the conversation history it chooses to include, and generation settings. The instruction is the part teams change most often and test least, and it is usually hard-coded as a string literal in a builder call.

This article explains how ADK Java turns an instruction into model input, including the exact placeholder rules (checked against the adk-java source at the time of writing), then builds the practices around it: prompts stored as versioned files, a small registry, logging which version produced which answer, tests that catch broken templates before deployment, and gradual rollout. Code that is ADK API is labelled as such; the registry and test code are patterns you write yourself.

Advertisement

What makes up an agent's prompt

The builder exposes these prompt-shaping settings:

  • instruction(String) or instruction(Instruction): the agent's core directions. A string is wrapped as Instruction.Static.
  • globalInstruction(...): text meant to apply across the agent tree, set on the root. The Python ADK documentation now recommends a plugin in place of its global_instruction parameter. In the Java source we checked, globalInstruction is present and not deprecated, so check the release you use before relying on it.
  • description(...): added to the agent's own system instruction by ADK's identity step, next to its name, and read by parent agents when deciding whether to delegate. A vague description costs tokens on every call and causes routing mistakes.
  • includeContents(IncludeContents.NONE): stops prior conversation history from being sent, which suits agents that should act only on state, such as a step in a pipeline.
  • outputKey("..."): stores the agent's final text in session state, which is how one agent's output becomes the next agent's {placeholder}.
  • generateContentConfig(...): temperature, output length and similar settings. These change behaviour as much as wording does, so version them with the prompt.

Instruction is a sealed interface with two records: Static(String instruction) and Provider(Function<ReadonlyContext, Single<String>>). At request time the agent resolves whichever one it has. The difference that matters most is that a static instruction has its placeholders filled from session state, and a provider's output is used verbatim.

Prompt filesprompts/support/v4.mdPrompt registryid, version, sha256Instruction.Staticplaceholders injectedInstruction.Provideryour code, no injectionSession stateapp:, user:, temp:, artifactsLlmRequestsystem instruction + contentsbeforeModelCallbacklog version, guardModelGemini or other BaseLlmload at boot{var}ctx.state()resolved textVersioned text lives outside the code;the agent only ever sees resolved strings.
How an instruction reaches the model in ADK Java. Static instructions have their placeholders filled from session state and artifacts; provider output is used exactly as returned. A callback sees the assembled request before the model call.

Placeholder rules in static instructions

In a static instruction, ADK finds brace-delimited placeholders and replaces them. The rules, from InstructionUtils.injectSessionState:

  • {name} is replaced with the session state value for name. Prefixed keys work too: {app:name}, {user:name} and {temp:name}, matching ADK's state scopes.
  • {artifact.report} is replaced with the text of the loaded artifact named report.
  • A trailing question mark makes a placeholder optional: {name?} becomes an empty string when the key is missing.
  • A required placeholder whose key is missing fails the call with IllegalArgumentException and the message Context variable not found. A missing required artifact fails similarly.
  • Text in braces that is not a valid identifier (with an optional recognised prefix) is left untouched. That is why a JSON example such as {"tier": "gold"} survives, since the quotes make it an invalid name.

The last two rules combine into the most common template bug. Suppose an instruction includes an example for the model: Reply like: Hello {customer_name}, .... To ADK that is a required placeholder, so if no state key customer_name exists, every call fails. If the key happens to exist, the example is silently filled with the current customer's name. Write examples with a different notation such as angle brackets, or produce them from a provider.

// ADK API: a static instruction with state placeholders.
LlmAgent support = LlmAgent.builder()
    .name("support_agent")
    .model("gemini-2.5-flash")
    .description("Answers billing and account questions for signed-in customers.")
    .instruction("""
        You are the support assistant for {app:product_name}.
        The customer is on the {user:tier} plan. Open tickets: {user:open_tickets?}
        Answer only from the policy below. If it does not cover the question, say so
        and offer to open a ticket.

        {artifact.refund_policy}
        """)
    .outputKey("support_answer")
    .build();
Advertisement

Instruction providers

A provider is a function from a ReadonlyContext to an RxJava Single<String>. It can read state through ctx.state(), look at ctx.userContent(), ctx.agentName() and ctx.events(), and call anything else it needs, asynchronously. Use it when the instruction's structure changes, not just its values: different sections for different tiers, a prompt version chosen per user, or text fetched from a registry.

Because provider output skips placeholder injection, a {user:tier} inside it reaches the model as literal text. That is a feature when your prompt contains braces of its own, and a trap when someone moves a static template into a provider and expects the placeholders to keep working.

// ADK API (Instruction.Provider) combined with your own registry.
Instruction supportInstruction = new Instruction.Provider(ctx -> {
    Map<String, Object> state = ctx.state();
    String tier = String.valueOf(state.getOrDefault("user:tier", "free"));
    PromptVersion pv = registry.select("support", ctx.userId());   // pinned or A/B
    String text = pv.render(Map.of(                                 // your renderer
        "tier", tier,
        "policy", policies.forTier(tier)));
    return Single.just(text);
});

LlmAgent support = LlmAgent.builder()
    .name("support_agent")
    .model("gemini-2.5-flash")
    .instruction(supportInstruction)
    .build();

Keep providers fast and deterministic for a given input. They run on every model call in the agent's turn, including calls made after tool results come back, so a slow remote lookup adds latency to every step. Cache registry lookups in memory and refresh them in the background.

Prompts as versioned files

A prompt embedded in Java source changes only on deploy, cannot be reviewed separately from the code, and gives no record of which wording produced a given answer. Treat prompts as versioned artifacts instead.

src/main/resources/prompts/
  support/
    v3.md          # immutable once released
    v4.md
    manifest.json  # {"default": "v3", "candidates": {"v4": 10}, "temperature": 0.2}
  triage/
    v1.md
// Your code, not ADK: load once at boot, hash, and refuse to start on errors.
public record PromptVersion(String id, String version, String sha256, String template) {
    public String render(Map<String, String> vars) {
        Matcher m = Pattern.compile("<<(\\w+)>>").matcher(template);
        StringBuilder out = new StringBuilder();
        while (m.find()) {
            String v = vars.get(m.group(1));
            if (v == null) throw new IllegalStateException(id + "/" + version + " needs " + m.group(1));
            m.appendReplacement(out, Matcher.quoteReplacement(v));
        }
        return m.appendTail(out).toString();
    }
}

public PromptVersion select(String id, String userId) {
    Manifest mf = manifests.get(id);
    int bucket = Math.floorMod(Hashing.murmur3_32_fixed().hashString(userId, UTF_8).asInt(), 100);
    int edge = 0;
    for (var c : mf.candidates().entrySet()) {          // stable: same user, same version
        edge += c.getValue();
        if (bucket < edge) return versions.get(id + "/" + c.getKey());
    }
    return versions.get(id + "/" + mf.defaultVersion());
}

Three choices in this sketch are deliberate. Released versions are immutable, so a version string always means the same text, and the content hash proves it. Your own templates use a different placeholder syntax (<<tier>>) from ADK's braces, so the two can never be confused. And bucketing hashes the user ID, so a user sees one version consistently instead of switching wording mid-conversation. The hashing call shown is Guava's; any stable hash works. Load prompts per environment with the configuration approach in environment configuration management.

Knowing which version answered

Every model answer should be traceable to the exact prompt version and hash. A provider only gets a read-only context, so it cannot write the chosen version into state. It does not need to: selection is a deterministic function of prompt ID and user ID, so a beforeModelCallback, whose signature is Maybe<LlmResponse> call(CallbackContext, LlmRequest.Builder), can recompute it and log it next to the agent name and invocation ID for every model call. Returning an empty Maybe lets the call proceed; returning a response short-circuits it, which is how guardrails and caches work (see ADK Java callbacks).

// ADK API: an observing callback that never changes the request.
LlmAgent support = LlmAgent.builder()
    .name("support_agent")
    .model("gemini-2.5-flash")
    .instruction(supportInstruction)
    .beforeModelCallback((callbackContext, request) -> {
        PromptVersion pv = registry.select("support", callbackContext.userId());
        log.info("model_call agent={} invocation={} prompt={}/{} sha={}",
            callbackContext.agentName(), callbackContext.invocationId(),
            pv.id(), pv.version(), pv.sha256());
        return Maybe.empty();
    })
    .build();

Join these log lines with your evaluation scores and user feedback, and you can answer the question that matters after every change: did version 4 do better than version 3, on which kinds of question, and at what cost in tokens? Where the request goes after the callback is described in model call orchestration.

Testing prompts in three layers

Prompt tests come in three layers, from cheap and deterministic to expensive and statistical.

  1. Template lint, on every build. Extract every ADK placeholder from every static instruction and check each required key against the keys your code writes (outputKey values, state writes in tools and callbacks, keys set at session creation). This catches the example-text bug and renamed state keys before anything runs. Also check that each prompt file renders with a full set of variables and fails with a missing one.
  2. Golden behaviour tests, on every pull request. A few dozen fixed conversations run against the real model at temperature 0 or low, with assertions on structure rather than exact wording: the agent refused, called the right tool, cited the policy, stayed under a length limit.
  3. Evaluation sets, before rollout. Hundreds of representative questions scored by rules and by a model grader, compared between the current and candidate versions. LLM-as-judge scoring covers building the grader, and ADK Java CI covers wiring the stages into a pipeline.
// JUnit, your code: fail the build when a required placeholder has no producer.
private static final Pattern ADK_VAR =
    Pattern.compile("\\{((?:app:|user:|temp:)?[A-Za-z_][A-Za-z0-9_]*)\\}");

@Test
void everyRequiredPlaceholderHasAProducer() {
    Set<String> produced = Set.of("app:product_name", "user:tier", "support_answer");
    Matcher m = ADK_VAR.matcher(SupportAgent.INSTRUCTION);
    while (m.find()) {
        assertTrue(produced.contains(m.group(1)), "no producer for {" + m.group(1) + "}");
    }
}

The regex deliberately ignores optional placeholders ending in a question mark and artifact references, which have their own checks. It approximates ADK's identifier rule; if a placeholder is too unusual for the lint to parse, that is a reason to simplify the placeholder.

Worked example: shipping v4 safely

Here is an illustrative change, with hypothetical numbers. A support agent runs prompt v3. Product wants gold-tier customers to be offered a callback. Someone edits the static instruction, adds the line Offer gold customers a call at {callback_number}, and deploys. Every conversation fails with Context variable not found, because no code writes callback_number. Under the process above, the change goes differently.

The new wording becomes support/v4.md and uses the renderer's <<callback_number>>, supplied by the provider from configuration. The lint passes, and a test confirms that v4 fails to render without callback_number. Golden tests show v4 offers the callback to gold users and never to free users. On a 400-question evaluation set, v4 matches v3 on accuracy and adds about 40 tokens per answer. The manifest sends 10 percent of users to v4. Logs carry the version on every call, and after a week the v4 cohort's resolution rate is no worse than v3's, so the default flips to v4. Rolling back is a one-line manifest change, and v3's file still exists.

Failure modes

  • Example text read as a placeholder. Curly braces in examples either crash the call or leak live state. Lint for them.
  • Providers that expect injection. Moving a template into a provider leaves its {keys} as literal text. Render inside the provider.
  • Edits to a released version. Changing v3 in place breaks every comparison made against v3. Make released files immutable, and check hashes at boot.
  • State scope mistakes. A value written under temp: does not persist beyond the current invocation, so the next turn's instruction cannot see it. Choose the prefix deliberately.
  • Global text in every agent. Global instructions are sent with every agent's requests, so long global text costs tokens on every call in the tree.
  • Settings drift. The wording is versioned but the temperature is set in code, so a version no longer fully describes behaviour. Keep generation settings in the manifest.

Trade-offs

Static versus provider. Static instructions are transparent and use ADK's state injection directly, but they fail hard on missing keys and cannot change structure. Providers allow any logic and remote sources, but they hide the final text from code review, so log or snapshot what they return.

Files in the jar versus a live registry. Prompts in resources deploy with the code, are reviewed in pull requests and cannot change without a release. A live registry allows instant changes and rollbacks without a deploy, but those changes bypass code review and CI, so give the registry its own approval step and tests.

Many small agents versus one long prompt. Splitting work across agents shortens each instruction and makes each one testable, at the cost of routing that depends on descriptions and more model calls.

What to do next

  1. List every LlmAgent and its instruction, global instruction, description and generation settings.
  2. Move instruction text into versioned resource files with a manifest, and make released versions immutable.
  3. Add a build-time lint that every required ADK placeholder has a producer and that examples contain no braces.
  4. Decide per agent between a static instruction and a provider, and render your own placeholders inside providers.
  5. Log prompt ID, version and hash on every model call from a beforeModelCallback.
  6. Write golden behaviour tests and a scored evaluation set, and run the evaluation on every candidate version.
  7. Roll out new versions to a stable hash bucket of users, compare outcomes, then flip the default or roll back.
Key takeaway: An ADK Java agent's prompt is its instruction plus global instruction, description, included history and generation settings. Static instructions get placeholders such as {user:tier}, {name?} and {artifact.x} filled from state, fail on missing required keys and leave invalid names alone; provider output is used verbatim. Store prompts as immutable versioned files with a manifest, use a placeholder syntax of your own inside providers, log the version on every model call, lint templates at build time, run golden tests and scored evaluations, and roll out each new version to a stable slice of users before making it the default.