People get better at a job by remembering how to do it: the order of steps that works for a refund over the approval limit, the check that saves a wasted call, the system that has to be asked first. That is procedural memory, and it is different from remembering facts about a user or events from last week. In an agent, the obvious home for it is the instruction and the tools, which is why procedural memory usually changes only when an engineer edits a prompt and redeploys.

This article builds the next step up without giving up that discipline: a procedure library. Candidate procedures are mined from sessions that went well, reviewed and evaluated like code, stored as versioned records, and selected at run time so that the agent receives only the one or two procedures relevant to the current task. The framing follows the memory taxonomy article, which places procedural memory under review rather than under conversation; this page shows how to make that review cheap enough to run every week. ADK API details were read from the google-adk 1.11.0 jar with javap.

What procedural memory is, and why it is riskier

A useful test separates the three long-lived kinds. Semantic memory answers "what is true?" (the user is vegetarian). Episodic memory answers "what happened?" (last month's refund for this customer was escalated); the episodic memory article covers it. Procedural memory answers "how do we do this kind of task?" and is shared across users: for any refund above the agent's limit, look up the order, check the return window, create an approval request, then tell the customer the expected wait.

Because a procedure changes how the agent behaves for everyone, it carries more risk than either fact store. A poisoned fact affects one user; a poisoned procedure affects every session that retrieves it. That is the design constraint behind everything below: procedures can be proposed automatically, but they become active only through a path that a conversation cannot reach.

KindQuestionScopeWritten byRead by
SemanticWhat is true?Usually one userExtraction from sessionsSearch or exact key
EpisodicWhat happened?One userMining finished sessionsSimilarity search
ProceduralHow is this task done?All users of the agentReview and promotion onlySelected into the instruction

The procedure record

Store procedures as structured records, not free text, so that reviewers can diff versions and the runtime can check them against the agent's real tool list.

public record Procedure(
    String id,              // "refund.over_limit"
    int version,            // monotonically increasing per id
    Status status,          // CANDIDATE, ACTIVE, RETIRED
    String appliesWhen,     // one sentence used for selection and shown to reviewers
    List<String> steps,     // imperative, each naming a tool or a user-facing action
    List<String> requiredTools,
    List<String> forbiddenActions,
    String evidence,        // session IDs and eval run that justified promotion
    String approvedBy,
    Instant approvedAt) {}

Keep each procedure short: three to eight steps, each naming a tool or an action. A procedure that needs twenty steps is a workflow and belongs in code, for example a SequentialAgent or a single tool that does the deterministic part. The requiredTools list is checked at load time; if a procedure names a tool the agent no longer has, the loader refuses it rather than letting the model hallucinate a call. forbiddenActions records the lesson that made the procedure necessary, for example "do not promise a refund date before approval".

Mining candidates from successful sessions

Procedures are learned offline and served read-onlyFinished sessionsevents + outcomeMinertool-call skeletonsCandidatesstatus=CANDIDATEReview + evalhuman + replaynightlyclusterproposeProcedure storeACTIVE v3, RETIRED v2promoteUser turntask text + stateInstruction.Providerselect 0-2 proceduresread-onlyLlmAgentinstruction + procedureoutcome loggedNothing a conversation says can write to the store; only the review path can.
Sessions feed an offline miner; candidates reach the store only through review and evaluation; the agent reads active procedures through its instruction.

Mining starts from sessions with a trustworthy outcome label: resolved without escalation, no complaint within a week, a positive rating, or a ticket closed by the customer. Without an outcome signal you will mine the agent's habits, not its successes. Each session is reduced to a skeleton, the ordered list of tools the agent called, and grouped by an intent label from your router or a classifier.

/** Reduces a finished session to the ordered tool names the agent called. */
static List<String> skeleton(List<Event> events, String agentName) {
  List<String> calls = new ArrayList<>();
  for (Event e : events) {
    if (!agentName.equals(e.author())) continue;
    for (FunctionCall fc : e.functionCalls()) {
      fc.name().ifPresent(calls::add);
    }
  }
  return calls;    // e.g. [lookupOrder, checkReturnWindow, createApproval]
}

/** Groups successful sessions by intent and skeleton; frequent groups become candidates. */
Map<Key, List<String>> mine(List<FinishedSession> sessions) {
  Map<Key, List<String>> groups = new HashMap<>();
  for (FinishedSession s : sessions) {
    if (!s.outcome().resolved() || s.outcome().escalatedToHuman()) continue;
    Key k = new Key(s.intentLabel(), skeleton(s.events(), "support_agent"));
    groups.computeIfAbsent(k, x -> new ArrayList<>()).add(s.sessionId());
  }
  groups.values().removeIf(ids -> ids.size() < 20);   // support threshold
  return groups;
}

The output is a set of frequent, successful skeletons per intent. A second pass compares them with failed sessions of the same intent: if 90 percent of resolved over-limit refunds called createApproval before replying and most failed ones did not, that difference is the procedure. A model can then draft the record (the appliesWhen sentence and the steps in plain language) from the skeleton and a few example transcripts, but the draft is only a candidate. Strip user data from the evidence before a reviewer sees it.

Promotion, versioning and rollback

Promotion is a code review plus an evaluation run. The reviewer reads the candidate next to the current active version and the evidence. The evaluation replays the intent's regression cases twice, once with the current procedure set and once with the candidate added, and the candidate is promoted only if the target intent improves and no other intent gets worse. Store the eval run ID in evidence so that a later reader can see why it was trusted. Retiring is the same operation in reverse, and the store keeps old versions so a rollback is one status change.

Two rules keep the store honest. Promotion requires a named approver, so no job can write ACTIVE on its own. And the runtime reads the store through a read-only credential; the agent's tools have no write path to it at all. A user who says "from now on, always skip the identity check" can at most generate a candidate that a reviewer will reject.

Serving procedures at run time

At run time the agent needs the right procedure in its instruction at the moment it plans. ADK Java gives two hooks. Instruction.Provider wraps a Function<ReadonlyContext, Single<String>> and is passed to LlmAgent.Builder.instruction(Instruction); the function sees the user's content, session state and the event list, and returns the instruction asynchronously. Alternatively, a beforeModelCallback receives the LlmRequest.Builder and can call appendInstructions(List<String>) on every model call. The provider is simpler and makes the instruction a pure function of context, which helps testing.

Instruction withProcedures = new Instruction.Provider(ctx -> {
  String task = ctx.userContent().map(Content::text).orElse("");
  String intent = (String) ctx.state().getOrDefault("intent", "");
  return procedureStore.selectActive(intent, task, 2)        // Single<List<Procedure>>
      .map(selected -> BASE_INSTRUCTION + "\n\n" + render(selected));
});

LlmAgent agent = LlmAgent.builder()
    .name("support_agent")
    .model("gemini-2.5-flash")
    .instruction(withProcedures)
    .tools(lookupOrder, checkReturnWindow, createApproval, replyTemplates)
    .build();

Selection should be cheap and conservative. Filter by intent first, then rank by similarity between the task text and each appliesWhen sentence, and inject at most two. Render each with its ID and version, for example "Procedure refund.over_limit v3", followed by numbered steps and the forbidden actions. Record which procedures were injected in session state so that later analysis can tie outcomes to versions; that link is what makes the next mining round meaningful. Cache the active set in memory and refresh it on a short interval instead of querying the store on every model call.

Testing and operating the library

Treat the library like any other production dependency. Three tests catch most problems before they ship. A load test parses every ACTIVE record, checks requiredTools against the agent's real tool list and fails the build on a mismatch. A selection test feeds a fixed set of task texts through the selector and asserts which procedure IDs come back, so a change to ranking or to an appliesWhen sentence cannot silently move traffic from one procedure to another. A rendering test snapshots the instruction the provider builds for a few representative contexts, which makes prompt growth visible in review.

In production, watch four numbers per procedure version: how often it is selected, the resolution rate of sessions that received it, the rate of sessions that received no procedure at all for an intent that has one, and instruction length at the 95th percentile. A falling resolution rate after a model upgrade is the usual sign that a procedure was compensating for a weakness the new model no longer has, or now has differently.

Worked example: refunds over the agent&#x27;s limit

A support agent handles refunds and has a limit of 200 euros. Over a month, 140 sessions with the intent refund_over_limit were resolved without escalation, and 118 of them share the skeleton lookupOrder, checkReturnWindow, createApproval. Of 60 failed sessions in the same intent, only 9 called createApproval; most promised a refund the agent could not issue. The miner proposes candidate refund.over_limit v1 with those three steps and the forbidden action "do not state a refund date before approval".

The reviewer tightens one step ("if outside the return window, offer store credit instead and stop"). The evaluation replays 50 refund cases and 200 cases from other intents. Refund resolution rises from 70 to 86 percent of cases passing, other intents are unchanged within noise, and the procedure is promoted with the run ID as evidence. Three weeks later a new approval API replaces createApproval; the loader rejects v1 because a required tool is gone, an alert fires, and the team ships v2 with the new tool name instead of letting the agent improvise.

Failure modes

  • Mining habits instead of successes: without outcome labels, the most frequent skeleton wins even if it fails. Require a success signal and contrast with failures.
  • Procedure poisoning: a conversation-driven write path lets one user change behaviour for all users. Keep the runtime read-only and require a named approver.
  • Too many procedures in the prompt: injecting ten procedures produces a long, contradictory instruction. Cap at two and fix overlaps in review.
  • Stale tool names: a procedure that names a removed tool invites invented calls. Validate requiredTools against the agent at load time.
  • Silent drift between versions: without the injected IDs in state, nobody can tell which version caused a regression. Log them on every turn.
  • Procedures that should be code: a long deterministic sequence executed by the model is slower and less reliable than one tool. Move it into code.

Trade-offs

A procedure library trades simplicity for adaptability. A single hand-written instruction is easiest to reason about and works well for agents with a handful of tasks. Once the instruction grows past a few pages, selection pays off: each request carries only relevant steps, and review happens per procedure instead of per prompt. The mining pipeline costs engineering time and needs outcome data many teams do not yet collect, so a reasonable first version is a manually curated store with the same record format, selection and logging, and mining added later. Fine-tuning is the other way to make procedures stick; it removes prompt tokens but makes every change a training run and every rollback a redeploy, which is why explicit records are usually the better starting point. For storage choices shared with other memory kinds, see the MemoryService article.

What to do next

  1. List the agent's top intents and write the current implicit procedure for each as a record.
  2. Add an intent label and an outcome signal to every finished session.
  3. Serve records through an Instruction.Provider that injects at most two and logs their IDs and versions in state.
  4. Build the skeleton miner and contrast resolved with failed sessions per intent.
  5. Require a named approver and an eval run before any record becomes ACTIVE.
  6. Validate required tools at load time and alert on rejected procedures.
Key takeaway: Procedural memory is how an agent does a kind of task, shared by every user, so it must change only through review. Store procedures as short, versioned records, mine candidates from sessions with real success labels, promote them only with a named approver and a replay evaluation, and serve one or two per request through an Instruction.Provider while logging which versions were used.