Anthropic says Claude leads 26% of measured AI R&D work
Anthropic says Claude led 26% of measured AI research and development work in August. Its internal index still found human supervision across every measured subset.
Claude now “leads” 26% of the artificial intelligence research and development work Anthropic measured inside the company, according to a prototype R&D Automation Index it published for August 2026. Anthropic defines “leads” as completing most of a task from a high-level prompt while a person supervises. It does not mean the system works autonomously.
More than 90% of the measured work reached at least the company’s “AI collaborates” level, where AI handles large portions of a task under close human direction. No measured subset reached the highest level, where AI operates fully autonomously with no human in the loop. The index covers Anthropic’s internal R&D process, not product releases such as Claude Code Projects for coordinated cloud agents.
Anthropic built the index from internal work records. For each week in July, it randomly sampled 20% of staff in every department involved in model R&D. A Claude research agent reviewed those employees’ Slack activity and internal documentation, generating about 15,000 granular tasks. Claude then organized the tasks into a frozen hierarchy of 542 nodes, including 378 leaf categories.
The company used sampled person-time to weight the categories by importance. Each employee received one unit of weight per sampled week, divided evenly among that person’s tasks. Anthropic called the method a “crude approximation.”
Claude agents researched how each category of work was performed. A separate Claude judge then assigned one of six automation levels using a scale proposed by Epoch AI. Anthropic said the judge and human raters chose the same level 59% of the time, compared with 35% exact agreement between pairs of human raters. Model and human ratings were within one level in 97% of comparisons. Anthropic also acknowledged room for disagreement at the boundary between collaboration and leadership.
The result is a company-run measurement, not an independent audit. Anthropic said using its own models to evaluate its systems could produce correlated errors, and that third-party or cross-developer verification would be needed to address that risk. The underlying Slack records, task-level evidence and weighting calculations were not published in the opened sources, so the 26% figure could not be independently reproduced.
The index also freezes its task basket from July 2026. Anthropic said a rising score shows that work in that basket is being automated, but does not by itself reveal whether employees have shifted to new kinds of work outside it.
More news

Sam Altman expected to attend Trump-Xi state dinner

Demo Stage premieres October 7. Tech Talks return October 15. Submit your project or talk proposal.
Dmytro Spodarets·Sep 22, 2026
Anthropic lays out Claude workflows and controls for financial advisors
