ReAct, from Yao and colleagues' 2022 paper, made a simple idea famous: let a language model alternate between reasoning about a task and taking actions with tools, feeding each observation back into the next round of reasoning. Nearly every agent framework since is a ReAct loop at heart. But the stack around the idea has changed. Models now emit structured tool calls instead of parsed text, can request several calls at once, and reason internally before they act. Context windows are large, prompt caching changes the cost curve, and many teams have learned that some of their agents should never have been agents.

This article revisits the pattern with those changes in mind. It assumes you know the basic loop; if not, start with ReAct explained, which builds the text version from scratch. Here we look at what to keep, what to drop, how much the loop really costs, how to constrain it without breaking caching, and how to decide with data when a fixed workflow should replace it.

Advertisement

What survived and what did not

The original format was text: the model wrote a line starting with Thought, then an Action line such as search[query], and the runtime stopped generation, ran the tool and appended an Observation line. Few-shot example trajectories taught the model the format. Almost every mechanical part of that has been replaced; the core idea has not.

2022 ReAct2026 practiceWhat it means for you
Text Thought / Action lines, regex parsingNative tool calls validated against JSON schemasParsing errors become schema-validation errors you can return as results
Stop sequence before ObservationGeneration ends at a tool call by constructionHallucinated observations largely disappear
One action per stepSeveral tool calls in one turnFewer turns for independent lookups; new ordering hazards
Visible thoughts as the reasoningInternal reasoning, often hidden or summarisedDo not ask for a verbose Thought on top; it duplicates cost
Few-shot trajectories in the promptTool descriptions and system instructionsTool docs are now the main lever on behaviour
Wikipedia API with three actionsDozens of tools, some with side effectsAction-space design and authorisation matter more than prompt wording

What survived is the essential insight: decisions grounded in fresh observations beat decisions made from memory, and the trajectory is an inspectable record of how the agent got its answer. The pattern's weaknesses survived too: it is myopic, deciding one turn at a time, and its cost grows with every turn.

The modern loop

Stable prefixsystem, tool schemas, taskGrowing historyturns, results, summariesModel turnreasoning + text + N tool callsExecutorphase allowlist, policy, approvalcallsParallel tool runtimetimeouts, truncation, idempotencyControllerbudget, phase, stop rulesappend resultsFinal answervalidated + tracedoneKeep the prefix byte-stable so it stays cached; enforce constraints in the executor, not by editing the tool list.
A modern ReAct loop: a stable cached prefix, a growing history, model turns that may carry several tool calls, an executor that enforces policy, and a controller that owns the budget and stop rules.

The loop below is provider-neutral: model is an adapter around whichever API you use that returns the turn's text, tool calls and token usage. Three design choices in it are the point of this article. Independent read-only calls run in parallel, but side-effecting calls run one at a time, in the order the model gave them. Constraints on which tools may be used are enforced in the executor, which returns an error result, rather than by changing the tool list. And the model's turn, including any reasoning items, is appended to the history exactly as returned.

import asyncio, json, time

class Budget:
    def __init__(self, max_turns=12, max_input_tokens=200_000, max_seconds=120):
        self.max_turns, self.max_in, self.deadline = max_turns, max_input_tokens, time.time() + max_seconds
        self.turns = self.input_tokens = 0
    def ok(self):
        return self.turns < self.max_turns and self.input_tokens < self.max_in and time.time() < self.deadline

PHASE_TOOLS = {                      # enforced here, never by changing the tool list
    "investigate": {"search", "read_file", "query_db", "declare_plan"},
    "act":         {"search", "read_file", "write_file", "open_ticket"},  # declare_plan is a real tool in tools/schemas
}
SIDE_EFFECTS = {"write_file", "open_ticket"}

async def run_call(call, tools, phase, approve, obs_chars=2000):
    name, args = call["name"], call["arguments"]
    if name not in PHASE_TOOLS[phase]:
        return f"Error: {name} is not allowed in phase '{phase}'. Allowed: {sorted(PHASE_TOOLS[phase])}"
    if name in SIDE_EFFECTS and not await approve(name, args):
        return f"Denied: a reviewer rejected {name}({json.dumps(args)[:200]})."
    try:
        out = await asyncio.wait_for(tools[name](**args), timeout=20)
    except Exception as e:                       # errors are observations, not crashes
        return f"Error: {type(e).__name__}: {str(e)[:300]}"
    s = str(out)
    return s if len(s) <= obs_chars else s[:obs_chars] + " ...[truncated; ask for a narrower range]"

async def react(model, tools, schemas, messages, approve, budget=None):
    budget = budget or Budget()
    phase = "investigate"
    while budget.ok():
        turn = await model(messages=messages, tools=schemas)   # provider adapter
        budget.turns += 1
        budget.input_tokens += turn.usage_input_tokens
        messages.append(turn.as_message())       # keep reasoning items exactly as returned
        if not turn.tool_calls:
            return turn.text, messages
        reads = [c for c in turn.tool_calls if c["name"] not in SIDE_EFFECTS]
        writes = [c for c in turn.tool_calls if c["name"] in SIDE_EFFECTS]
        results = await asyncio.gather(*(run_call(c, tools, phase, approve) for c in reads))
        for c in writes:                           # side effects run one at a time, in order
            results.append(await run_call(c, tools, phase, approve))
        for c, r in zip(reads + writes, results):
            messages.append({"role": "tool", "tool_call_id": c["id"], "content": r})
        if any(c["name"] == "declare_plan" for c in turn.tool_calls):
            phase = "act"
    return None, messages                        # out of budget: caller decides what to do

The phase switch is deliberately crude: the agent must call a declare_plan tool before it may write anything. That one rule turns an unconstrained loop into a two-state machine with an audit point between investigating and acting, a pattern developed further in state machines for agents.

Advertisement

Worked example: what a twelve-turn loop costs

Every turn resends the whole context, so input tokens grow roughly with the square of the number of turns. Put numbers on it. Suppose the stable prefix (system prompt, tool schemas, task) is 3,000 tokens, and each turn adds 150 tokens of model output plus 800 tokens of tool results, so 950 tokens per turn. Turn k reads 3,000 + 950 x (k - 1) input tokens.

Over 12 turns the input total is 12 x 3,000 + 950 x (0 + 1 + ... + 11) = 36,000 + 950 x 66 = 98,700 input tokens, against only 1,800 output tokens, plus whatever internal reasoning the model generates. Input dominates, and most of it is the same tokens re-read.

ChangeInput tokens over the runVersus baseline
Baseline: 12 turns, 800-token results98,700-
Truncate results to 300 tokens (450 per turn)36,000 + 450 x 66 = 65,700-33%
Parallel calls: 6 turns, twice the results per turn (150 + 1,600)18,000 + 1,750 x 15 = 44,250-55%
Both: 6 turns at 750 per turn18,000 + 750 x 15 = 29,250-70%

Halving the number of turns more than halves the bill, because the quadratic term shrinks faster than the per-turn content grows. That is the real economic case for parallel tool calls and for coarser tools that answer a question in one call instead of three.

Prompt caching changes the picture again. Providers that cache a repeated prefix charge less, and respond faster, for the cached portion; in a loop, every turn's prefix is the previous turn's entire context. The catch is that caching works only on byte-identical prefixes. Rewriting history (summarising old turns, reordering, editing the tool list, putting a timestamp in the system prompt) invalidates everything after the change. The practical rules: keep the prefix stable, append rather than edit, and when you must compact, do it rarely and in large chunks, as discussed in context compaction. Check your provider's documentation for its cache rules and prices rather than assuming them.

Reasoning models inside the loop

Models that reason internally before answering change ReAct in three ways. First, the visible Thought line is redundant. Asking a reasoning model to also write out its thinking costs tokens and adds a second, possibly inconsistent, account of the same decision. Ask instead for a one-line statement of intent alongside each tool call, which is useful in logs.

Second, reasoning now happens between tool calls as well as before the first. Some APIs return reasoning items that must be passed back unchanged with the tool results on the next turn; dropping or editing them can cause an error or degrade the next decision. That is why the loop above appends the turn exactly as returned. Read your provider's tool-use documentation on this point.

Third, per-turn latency and cost rise with reasoning effort, which shifts the balance towards fewer, better-chosen turns. A reasoning model will often plan several steps internally and then issue a batch of parallel calls, which is a plan-and-execute pattern emerging inside a ReAct loop. Let it: count turns, not tool calls, when you set budgets.

One caution carries over from the original paper's era: a stated rationale, visible or summarised, is not guaranteed to be the real cause of an action. Use stated intents for debugging, but audit agents on what they did, the tool calls and their arguments, not on what they said they were thinking.

Constraining the action space

The biggest reliability gains rarely come from prompt wording. They come from the shape of the action space. Coarse, purpose-built tools (get_order_status(order_id)) beat fine-grained general ones (query_db(sql)) on both cost and error rate, because each call does more and offers fewer ways to be wrong. Tools with side effects should be idempotent, carry a request id, and sit behind a policy check or human approval.

Phase-dependent rules belong in the executor. It is tempting to send a different tool list each turn, exposing only what the current phase allows, but tool schemas normally sit near the start of the prompt, so changing them busts the cache on every turn and confuses the model about what exists. A stable list plus a clear error result ("write_file is not allowed in phase 'investigate'") keeps the cache warm and teaches the model the rule in context.

When a workflow should replace the loop

ReAct is the right default when the next step genuinely depends on what was just observed: debugging, open-ended research, exploring an unfamiliar system. It is the wrong tool when the path is predictable, because you pay for a model decision at every turn to rediscover a fixed sequence. The way to tell is to look at the traces. If most successful runs call the same tools in the same order, encode that order as a workflow with model calls only at the points that need judgment, and keep ReAct as the fallback for cases the workflow cannot handle.

Make the decision with a replay harness over a fixed task set, not by intuition. Run each variant several times per task, because agents are stochastic, and compare success, turns, tokens and tail latency:

def compare(variants, tasks, runs_per_task=3):
    """variants: name -> callable(task) returning dict(success, turns, input_tokens, seconds)."""
    report = {}
    for name, run in variants.items():
        rows = [run(t) for t in tasks for _ in range(runs_per_task)]
        rows.sort(key=lambda r: r["seconds"])
        report[name] = dict(
            success=sum(r["success"] for r in rows) / len(rows),
            mean_turns=sum(r["turns"] for r in rows) / len(rows),
            mean_input_tokens=sum(r["input_tokens"] for r in rows) / len(rows),
            p95_seconds=rows[int(0.95 * (len(rows) - 1))]["seconds"],
        )
    return report

# variants = {"react": run_react, "plan_then_react": run_planned, "workflow": run_fixed_pipeline}
PatternPrefer it whenWatch out for
Plain ReActNext step depends on the last observationTurn count and quadratic cost
Plan, then ReAct per stepTasks decompose into known stagesStale plans; allow replanning on failure
Fixed workflow with model nodesTraces show one dominant pathRigidity; keep a ReAct fallback
Planned batch of calls (ReWOO style)Independent lookups, cost-sensitiveNo adaptation to surprising results

For how to make the success column trustworthy, see agent evaluators.

New failure modes

  • Conflicting parallel writes. Two side-effecting calls in one turn race or run in the wrong order. Serialise writes and make them idempotent.
  • Cache thrash. Costs jump after a harmless-looking change, such as a timestamp in the system prompt or reordered tool schemas. Keep the prefix byte-stable and monitor the cached-token ratio.
  • Dropped reasoning items. A middleware layer strips unknown message parts, and tool-use turns start failing or degrading. Pass turns through untouched.
  • Overthinking per turn. High reasoning effort on every trivial lookup doubles latency. Tune effort per task type.
  • Budget exhaustion without an answer. The loop ends with nothing useful. Return a partial answer with what was found and what remains, and count these runs as failures in evaluation.
  • Injected instructions in tool results. Still the most serious risk: results are untrusted data, and side effects must be gated regardless of what a result says.

What to do next

  1. Measure your agent's turns per task and input tokens per run, and compute the quadratic cost for your real numbers.
  2. Enable parallel read-only tool calls and serialise side-effecting ones; truncate tool results with a marker.
  3. Freeze the prompt prefix and tool list, move phase rules into the executor, and track the cached-token ratio.
  4. Pass reasoning items back exactly as returned and drop any instruction to write visible thoughts.
  5. Replace fine-grained tools with coarse purpose-built ones where traces show repeated multi-call sequences.
  6. Build a replay harness, compare ReAct with a plan-first and a workflow variant, and keep whichever wins on success per unit cost.
Key takeaway: ReAct's insight, grounding each decision in a fresh observation, is as sound as ever, but its mechanics have moved: native and parallel tool calls, internal reasoning, cached prefixes. Count turns because cost is quadratic in them, keep the prefix stable and enforce rules in the executor, pass reasoning back untouched, gate side effects, and let a replay harness decide which tasks should leave the loop for a workflow.