The news

Primary: Anthropic, "Measurements for understanding the pace of AI development inside frontier labs".

Three measurements:

  1. AI-led AI R&D (Automation Index). August 2026: 26% of Anthropic AI R&D work at AL4 ("leads"); >90% at or above AL3 ("collaborates"); no measured AL5 autonomy.
  2. Oversight of agents. Snapshot of the most-used internal platform: ~30,000 agents; online monitors before execution and offline monitors after; coverage stated at 100% for both; August sample of over a billion online decisions with ~0.002% blocked.
  3. Compute allocation. Week of 13-20 July: ~6% of AI R&D compute and ~12% of AI-driven AI R&D compute tagged to safety, with conservative coding of dual-purpose work as non-safety.

Anthropic says it plans to embed independent third-party evaluators with internal-like access to verify practices and metrics.

Why it matters

Pacing debates need numbers labs can publish without waiting for a regulator's template. These three are Anthropic's opening bid for comparable reporting.

What to do now

  • Policy and eval teams: read the Appendix before citing 26%. The figure is person-time-weighted AL4 on a frozen task tree.
  • Other labs: Anthropic explicitly invites parallel publication with shared methodology and third-party checks.
  • Reporters: keep oversight and compute figures in separate sentences from the 26% AL4 headline.

Caveats

Cross-lab comparability is UNKNOWN until others publish on the same AL definitions. Judge models rating their own lab's work remain a stated methodological risk.