The news
Primary: Anthropic, "Measurements for understanding the pace of AI development inside frontier labs".
Three measurements:
- AI-led AI R&D (Automation Index). August 2026: 26% of Anthropic AI R&D work at AL4 ("leads"); >90% at or above AL3 ("collaborates"); no measured AL5 autonomy.
- Oversight of agents. Snapshot of the most-used internal platform: ~30,000 agents; online monitors before execution and offline monitors after; coverage stated at 100% for both; August sample of over a billion online decisions with ~0.002% blocked.
- Compute allocation. Week of 13-20 July: ~6% of AI R&D compute and ~12% of AI-driven AI R&D compute tagged to safety, with conservative coding of dual-purpose work as non-safety.
Anthropic says it plans to embed independent third-party evaluators with internal-like access to verify practices and metrics.
Why it matters
Pacing debates need numbers labs can publish without waiting for a regulator's template. These three are Anthropic's opening bid for comparable reporting.
What to do now
- Policy and eval teams: read the Appendix before citing 26%. The figure is person-time-weighted AL4 on a frozen task tree.
- Other labs: Anthropic explicitly invites parallel publication with shared methodology and third-party checks.
- Reporters: keep oversight and compute figures in separate sentences from the 26% AL4 headline.
Caveats
Cross-lab comparability is UNKNOWN until others publish on the same AL definitions. Judge models rating their own lab's work remain a stated methodological risk.
