News

Anthropic says Claude leads 26% of measured AI R&D work

Anthropic says Claude led 26% of measured AI research and development work in August. Its internal index still found human supervision across every measured subset.

D
Sep 22, 2026 · 2 min read

Claude now “leads” 26% of the artificial intelligence research and development work Anthropic measured inside the company, according to a prototype R&D Automation Index it published for August 2026. Anthropic defines “leads” as completing most of a task from a high-level prompt while a person supervises. It does not mean the system works autonomously.

More than 90% of the measured work reached at least the company’s “AI collaborates” level, where AI handles large portions of a task under close human direction. No measured subset reached the highest level, where AI operates fully autonomously with no human in the loop. The index covers Anthropic’s internal R&D process, not product releases such as Claude Code Projects for coordinated cloud agents.

Anthropic built the index from internal work records. For each week in July, it randomly sampled 20% of staff in every department involved in model R&D. A Claude research agent reviewed those employees’ Slack activity and internal documentation, generating about 15,000 granular tasks. Claude then organized the tasks into a frozen hierarchy of 542 nodes, including 378 leaf categories.

The company used sampled person-time to weight the categories by importance. Each employee received one unit of weight per sampled week, divided evenly among that person’s tasks. Anthropic called the method a “crude approximation.”

Claude agents researched how each category of work was performed. A separate Claude judge then assigned one of six automation levels using a scale proposed by Epoch AI. Anthropic said the judge and human raters chose the same level 59% of the time, compared with 35% exact agreement between pairs of human raters. Model and human ratings were within one level in 97% of comparisons. Anthropic also acknowledged room for disagreement at the boundary between collaboration and leadership.

The result is a company-run measurement, not an independent audit. Anthropic said using its own models to evaluate its systems could produce correlated errors, and that third-party or cross-developer verification would be needed to address that risk. The underlying Slack records, task-level evidence and weighting calculations were not published in the opened sources, so the 26% figure could not be independently reproduced.

The index also freezes its task basket from July 2026. Anthropic said a rising score shows that work in that basket is being automated, but does not by itself reveal whether employees have shifted to new kinds of work outside it.

More news