Document

Metric 1 - AI-led AI R&D

ItemValue
InstrumentAnthropic R&D Automation Index
ScaleEpoch AI Automation Levels AL0-AL5
As ofAugust 2026
AL4 ("leads") share26% of Anthropic AI R&D work
≥ AL3 ("collaborates")>90%
AL5 fully autonomousnone in measured subsets

Appendix method: sample staff weeks in July 2026 → ~15,000 tasks → frozen tree (542 nodes / 378 leaves) → Claude research + Claude judge ratings → person-time weights. Human agreement check reported in-appendix.

Metric 2 - Oversight of AI agents

ItemOnline monitorsOffline monitors
RoleBlock/steer before executionReview after the fact
Coverage (stated)100% of actions on the measured platform100% ingested after the fact
Platform scale~30,000 agents on most-used internal research/engineering platformsame
August signal>1B decisions; ~0.002% blocked (~1 in 47,000)~1-2 / 1000 transcripts flagged; ~50/week high-priority human escalations

Metric 3 - Compute allocation

ItemValue
Window13-20 July (one week snapshot)
Safety share of AI R&D compute~6%
Safety share of AI-driven AI R&D compute~12%
Coding ruleDual-purpose tokens counted as non-safety (conservative)
NoteSafeguards-classifier compute excluded from these shares

Appendix pointers

  • AL3 vs AL4 concrete pipeline-fix vignette in footnotes
  • Compute classifier prompt excerpt for safety vs R&D labeling
  • Plan to embed third-party evaluators with internal-comparable access

Reading rule

Cite numbers with as-of dates and definitions. The 26% figure is AL4 on a person-time-weighted AI R&D basket. It is not headcount, revenue, or total company work.

Source

  1. Institute post