Anthropic · 2026-09-17 · major
Anthropic publishes pace metrics — Claude leads 26% of its own AI research
Anthropic proposes three measurements frontier labs could publish about their own AI development and reports its August 2026 numbers: Claude leads 26% of Anthropic's AI R&D work, and about 6% of AI R&D compute goes to safety.

Three numbers Anthropic thinks every frontier lab should publish, starting with its own.
Key specs
| Ai led r&d | 26% |
|---|---|
| Safety compute | 6% |
Quick facts
| Publisher | Anthropic Institute |
|---|---|
| Measurement 1 | AI-Led AI R&D (R&D Automation Index) |
| Measurement 2 | Oversight of AI agents |
| Measurement 3 | Compute allocated to safety |
| Data period | August 2026 |
| Agent monitoring | ~30,000 concurrent agents, 100% of actions monitored |
| Actions blocked | 0.002%, about 1 in 47,000 |
What is it?
Three proposed measurements make up this post from the Anthropic Institute, each aimed at making frontier-lab activity visible from outside. The first tracks how much AI research is done by AI. The second covers how agent actions are supervised. The third splits compute between safety work and everything else. Anthropic publishes its own August 2026 figures against all three rather than only proposing them.
How does it work?
The R&D Automation Index is the load-bearing piece. Anthropic grades work on an AL0-AL5 automation scale borrowed from Epoch AI, where each level describes how much a human still contributes. On that scale Claude 'leads' 26% of the company's AI R&D, more than 90% of the work sits at or above 'AI collaborates', and no measured R&D work is fully autonomous. The oversight metric reports coverage, review latency and escalation rates from the monitors that sit in front of agent actions.
Why does it matter?
Arguments about how fast AI is moving usually run on anecdote. Putting a number on how much of a frontier lab's own research its models already do turns that into something regulators, evaluators and rival labs can check and compare. Anthropic pairs the metrics with a plan to embed independent third-party evaluators from several organisations inside the company, which is what would make the figures verifiable rather than self-reported.
Who is it for?
AI policy researchers and safety evaluators
Frequently asked questions
- What does it mean that Claude 'leads' 26% of Anthropic's AI R&D?
- Anthropic grades each piece of research work on the AL0-AL5 automation scale from Epoch AI. 'Leads' is a specific rung on that scale, above 'AI collaborates' but below full autonomy. Anthropic reports that no measured R&D work reaches the fully autonomous level, so the 26% figure describes Claude directing work that humans still review, not unsupervised research.
- Are other AI labs going to publish these measurements too?
- Anthropic frames the three measurements as a proposal for the industry, not an agreement any other lab has signed. The post argues that shared metrics would give the public visibility into frontier development, and Anthropic publishes its own figures first. Whether OpenAI, Google DeepMind or others adopt the same scale is unresolved in the post.
- How does this connect to Dario Amodei's 'We Must Pace the Frontier'?
- The Anthropic Institute post is the operational follow-up to that essay, which it cites as the motivation. Where Dario Amodei argued that labs should slow how fast capabilities grow, this post proposes the specific indicators that would show whether they are, and repeats the commitment to embed independent third-party evaluators inside Anthropic.
- Can outsiders check the underlying data?
- Partly. Anthropic links a redacted August 2026 Risk Report and its Advanced AI Framework document alongside the post, so the reasoning is published rather than just the headline figures. The numbers themselves remain self-measured for now; the embedded third-party evaluators Anthropic describes are the mechanism intended to make them independently verifiable.
Try it
https://www.anthropic.com/institute/measuring-pace-of-ai-development