Asana says browser-agent redesign cut test costs 76x
Asana says a redesigned browser-agent workflow averaged $0.47 and about four minutes per test, while its controlled study shows why the headline comparison needs qualification.
Asana said its redesigned browser-agent workflow averaged $0.47 in estimated model cost and roughly four minutes per run with GPT-6.1 Sol. That made it 76 times cheaper and five times faster than the company’s original production setup on an anonymized Model B. Asana said the changes were subsequently shipped to StackAI.
The headline result combines a workflow redesign with a model change; it is not a before-and-after comparison for GPT-6.1 Sol alone. OpenAI said the original Model B runs cost at least $36.21 and took at least 22.5 minutes. Some runs hit the step limit before finishing, making those averages—and the resulting fold reductions—lower bounds. Within GPT-6.1 Sol at the larger history budget, the new caching and screenshot policy cut estimated cost from $1.97 to $0.47 per run, a fourfold change. On Model B, the optimized workflow cut estimated cost to $1.24 per run, a 29-fold reduction from the original setup, and ran four times faster.
The redesign cached the agent’s browsing history, raised its history budget from 120,000 to 480,000 characters, and moved screenshot cleanup to batch pruning. The system could accumulate as many as 20 screenshots before retaining only the newest one. The earlier agent cached fixed instructions and tool definitions, but continually deleted screenshots and trimmed old page text. Those edits changed the request prefix on almost every step and prevented repeated reuse of cached input.
Asana and OpenAI said 89% of GPT-6.1 Sol input in the optimized condition came from cache. OpenAI said cache reads for the model were priced at 5% of uncached input, making each call about three times cheaper. But StackAI by Asana said caching alone was not enough. At the larger history budget, caching without batch pruning cost more than not caching on three of the four tested models because earlier history kept being rewritten. Its engineering lesson: keep history append-only and, when pruning is necessary, do it in large batches.
The company tested six caching and history policies at two character budgets across GPT-6.1 Sol and three anonymized frontier models. Three runs per condition yielded 144 runs, followed by 12 more. Each condition asked the agent to collect six fields for each of 32 books from a public demonstration catalog, for 192 expected facts. OpenAI said every optimized-workflow run finished and returned the correct answer. Asana said every best-condition run encountered all 192 expected facts.
The evidence remains a company-run study, not an independent audit. The opened sources did not include raw traces, complete per-run values, or complete standard deviations, while the anonymized models prevent outside replication of the cross-model comparison. Asana said three or four runs per condition could reveal broad patterns but could not reliably distinguish differences of only a few percent. The benchmark covered one catalog-extraction task; no run reached the 480,000-character cap, and compaction was outside the study. OpenAI also said some chart markers and error bars were approximate reconstructions, and Model B was tested in a different phase from Model C and GPT-6.1 Sol.
Asana said GPT-6 Astra in Codex helped audit and instrument the code, refactor the test harness, launch runs, and analyze traces. People set the goals and standards and reviewed the conclusions. The company then turned the findings into tickets and reviewed pull requests before shipping the browser-navigation changes.
More news

Jump Trading Uses GPT-6 Astra for Days-Long Quant Research

California expands behavioral-health county profile built on Databricks

Harvard study ties coding-agent gains to heavier code review
