Anthropic has published the first results from a prototype R&D Automation Index, reporting that its Claude models now "lead" 26% of the company's own AI research and development work as of August 2026 ā up from under 1% in February. The measurement, released on 17 September in an Anthropic Institute post authored by Marina Favaro and Phillie Wright, is the most concrete public number any frontier lab has put on how much of its own research an AI system is running.
What "Leads" Means ā And Does Not Mean
The scores use an Automation Level scale developed by Epoch AI, running from AL0 (no AI involvement) to AL5 (fully autonomous with no human in the loop). The 26% figure refers specifically to AL4, where the model completes most of a task end to end from a high-level prompt while a human supervises.
The distribution is as informative as the headline. None of the measured work sits at AL5. More than 90% sits at AL3 or above, meaning AI involvement is now near-universal across Anthropic's research tasks even where a human still drives. The company was explicit that the number does not mean Claude is independently building its successor ā no measured category shows the model operating fully autonomously.
The Methodology Is the Interesting Part
Anthropic built the index from internal work records, including Slack messages and company documentation. During each week of July 2026, it randomly sampled 20% of staff from departments involved in model R&D, producing roughly 15,000 individual tasks.
Claude then organised those tasks into a hierarchy of 542 categories, 378 of which are leaves ā granular units such as "eval platform defect diagnosis and fixes." The tree is deliberately frozen, so every future measurement runs against the same fixed basket of work rather than a drifting definition. For each node, a Claude agent researches how that work is actually performed and an independent Claude judge assigns an automation level, restricted to evidence from the month being rated or earlier. Tasks are weighted by person-time.
The frozen-basket design is what makes the index a time series rather than a snapshot. It also imports a known weakness: a fixed taxonomy built in July 2026 will slowly stop describing the work as the work changes shape.
Two More Numbers
Anthropic paired the index with two other self-reported metrics, and argued that any frontier developer could publish the same three:
- Agent oversight. Roughly 30,000 internal agents run concurrently, 100% of them monitored, with approximately one action in 47,000 blocked by monitors.
- Compute allocation to safety. About 6% of total compute, or 12% of the AI-driven share, measured over a single week from 13 to 20 July 2026.
The blocking rate deserves careful reading. A rate of 0.002% is equally consistent with agents rarely attempting anything that warrants intervention and with a monitor tuned to intervene rarely by design. Without knowing the monitor's precision and recall, the number describes the monitor as much as it describes the agents.
Why It Matters
Recursive self-improvement has been the central abstraction in AI risk discussion for over a decade, and it has been almost entirely unmeasured. What Anthropic has done is convert a speculative concept into an instrument with a defined scale, a fixed denominator, and a publication cadence ā the difference between arguing about whether the curve is steep and being able to point at it.
The trajectory is the story. Going from under 1% to 26% in six months is a gradient that, if it holds even approximately, puts the question of AL5 work on a near-term timetable rather than a philosophical one. Anthropic itself flags that the February baseline is an upper bound rather than a point estimate, which softens the slope somewhat but not the direction.
There is also a competitive dimension the lab did not emphasise. If a quarter of a frontier lab's research is model-led, the cost structure of frontier research changes ā and labs with better internal tooling compound faster than labs with better researchers. That is a different competitive landscape than the one the industry has been optimising for.
The Verification Gap
Every figure in the disclosure is Anthropic measuring itself. No external party has verified any of the ratings.
This is the honest limit of the exercise, and it sits awkwardly beside the publication's own context. The post was the lab's first substantive release since Dario Amodei's 12 September essay urging the industry to pace the frontier and committing Anthropic to embedded third-party evaluators. An index that invites comparison across labs while being unverifiable at any of them is a first draft, not a standard.
Still, first drafts are how standards start. The proposal ā three numbers, published on a cadence, comparable over time and across organisations ā is cheap for any lab to adopt and hard to argue against in principle. Bloomberg and the Washington Post both covered the release on 17 September, and the pressure now falls on OpenAI and Google DeepMind to either publish equivalents or explain why they will not.
