Quoted claim
Headline form: 26% of Anthropic R&D is AI-led. Sometimes shortened further to "AI builds a quarter of Anthropic."
Where made
Primary: Anthropic, "Measurements for understanding the pace of AI development inside frontier labs". Section (1) Measuring AI-led AI R&D introduces the Anthropic R&D Automation Index.
Quoted finding as of August 2026:
- Claude is not operating fully autonomously for any measured subset of AI R&D work
- Claude "leads" 26% of Anthropic's AI R&D work
- Share of work at or above "AI collaborates" is above 90%
AL scale (Epoch AI, as adopted): AL0 none through AL5 fully autonomous. AL4 ("leads"): AI can complete most of the task end-to-end from a high-level prompt while a human supervises. AL3 ("collaborates"): large chunks under close human direction.
Supporting number / method
Appendix method, compressed:
- Build a bottom-up task list from staff samples and internal work records (Slack and docs), ~15,000 granular tasks for July 2026 weeks
- Organize into a frozen hierarchical tree (542 nodes, 378 leaves)
- Rate each node with a Claude research agent plus an independent Claude judge on the AL scale, using only evidence from the rated month or earlier
- Weight nodes by sampled person-time
Human cross-check: model-versus-human exact agreement 59%; within one level 97%. Humans agreed with each other at 35% exact on the same exercise.
Companion metrics in the same post (not the 26% claim): ~30,000 agents on the most-used internal platform; online and offline monitor coverage claimed at 100% of those agents' actions; one-week compute snapshot with ~6% of AI R&D compute and ~12% of AI-driven AI R&D compute tagged to safety under a conservative classifier.
What would have to be true
For "26% of Anthropic R&D is AI-led" without qualifiers to stay accurate, readers would need the AL4 definition, the AI R&D basket (not all company functions), the August 2026 as-of date, and the person-time weighting. Dropping those turns a methodology-bound index into a fake headcount or autonomy claim. AL5 (full autonomy) is explicitly not reached in the measured subsets.
Verdict
supported when stated as: as of August 2026, Anthropic's Automation Index rates Claude at AL4 ("leads") on 26% of its AI R&D task basket under the published methodology.
overstated as 26% of employees replaced, 26% of compute, company-wide R&D including non-AI work, or proof that recursive self-improvement has arrived.
Language we will use instead
We say Anthropic reports Claude leading 26% of AI R&D tasks at AL4 on its August 2026 index. We keep AL definitions and the frozen task basket in the sentence when the number is used as evidence of pace.
