Asana's post, "How we cut a browser agent's cost 76x and made it 5x faster by keeping its cache intact," is dated October 8, 2026; OpenAI published its own case study the next day. The agent belongs to StackAI, the no-code platform Asana acquired, where customers build workflows that navigate websites, fill out forms and gather information.
What was going wrong
A browser agent "resends its tools, system prompt and growing history of page text and screenshots on every call." Prompt caching makes repeated input cheap – cache reads cost 0.05x to 0.1x the standard input price on the models tested – "But a cache only reuses the longest unchanged prefix of a request." StackAI's agent cached its tools and prompt but not its history, and it removed the previous screenshot at every step and trimmed older text to fit a 120,000-character budget. "Each edit altered an earlier part of the request, so reuse broke on almost every call."
The fix
- Cache the history, with a cache marker on the latest tool result.
- Prune in batches. At 20:1 the agent keeps up to 20 screenshots, then cuts back to one, "so about 19 consecutive calls reuse the cached history."
- Raise the budget from 120,000 to 480,000 characters so old text is no longer trimmed.
Results
Six caching and history policies at two budgets were tested on GPT-6.1 Sol and three unnamed frontier models, three runs each – "144 runs, plus a 12-run follow-up." The task was collecting six fields for each of 32 books from a public demo catalogue. On the original production model the best setup cut cost per run 29x and ran 4x faster; on GPT-6.1 Sol it cost 76x less than the original production setup, about $0.47 a run against at least $36.21, "reading 89% of its input from cache." Two findings generalise:
- "Caching alone is not enough." Without batch pruning, caching the history at the larger budget cost more than not caching on three of four models, because "the cache was continually rewritten and rarely read."
- "The budget has to fit the model." At 120,000 characters, the newest unnamed model answered in none of 18 runs and GPT-6.1 Sol in 3; at 480,000, every run on both answered.
How the work was done
GPT-6 Astra in Codex "audited the code, instrumented every request," refactored it to run workflows in parallel, launched the runs and analysed traces; "Humans set the goal and the standards and reviewed the conclusions." Every session was recorded in Asana's Command delivery platform, which the agents reached through MCP, and findings became tickets and reviewed pull requests. StackAI CTO Frank Hidalgo told OpenAI that work he estimated at one to two months by hand took about a week: "I'd set a /goal before going to bed and review the results in the morning." His conclusion: "Shipping speed is no longer the bottleneck; human attention is."
What remains uncertain. Both write-ups come from the companies involved; the competing models are unnamed; and Asana itself notes that with three or four runs per condition the study "shows broad patterns rather than distinguishing conditions only a few percent apart."
Analysis: the practical checklist is cheap to apply to any long-running agent: keep the history append-only, prune in large batches, size the budget to the model, and measure cache reads with the provider's own counters. Many agent cost problems are cache-breaking problems. The human-attention point is the adoption constraint to plan for: once an agent can run experiments overnight, review capacity, not compute, sets the pace.
Asana's StackAI agent cached its prompt but broke cache reuse on nearly every call by trimming its history; caching the history, pruning screenshots in batches and raising the budget cut cost 76x on GPT-6.1 Sol in a 144-run study that GPT-6 Astra in Codex largely ran while humans set goals and reviewed results.