A technician kneeling beside open server racks in a data center, adjusting cabled computing hardware.
← Back to Blog
Agentic AI

Claude now leads 26% of Anthropic's own AI research and development

Head portrait of Alex Goryachev
Alex Goryachev·September 20, 2026·5 min read

In 6 months, the share of Anthropic's AI research led by Claude went from under 1% to 26%, and the company published that number rather than discovering it in hindsight.

Key Takeways

  • Claude led 26% of Anthropic's AI research and development work in August 2026, up from under 1% in February 2026, according to research published by the Anthropic Institute.
  • Anthropic defines "leads" as the AI completing most of a task end to end from a high-level prompt while a person supervises, and states that Claude is not operating fully autonomously for any measured subset of its AI R&D work.
  • Anthropic's online monitors blocked about 0.002% of agent decisions, roughly 1 in 47,000, across approximately a billion agent decisions analyzed in August 2026.
  • Anthropic proposed 3 new measures going forward: an R&D Automation Index, Agent Oversight Metrics, and a Compute Allocation Measurement tracking how much computing power goes to safety work.

In February 2026, Claude led under 1% of the AI research and development work inside Anthropic. By August, it led 26%. Same company, same category of work, 6 months apart.

Anthropic published that figure about itself, in research from its Institute reported on September 17. Nobody at the company was laid off in this story and nobody was replaced. What changed is the shape of the work, and it changed faster than most careers are built to change.

The people on the other side of that number are AI researchers at a frontier lab. If any job looked safe from automation 3 years ago, it was that one.

What 26% actually measures

"Leads" is a specific rating on Anthropic's own automation scale, which draws on the kind of scale Epoch AI uses. It describes work where the model completes most of a task end to end from a high-level prompt while a person supervises the result. That is not autonomy, and Anthropic says so directly: Claude is "not operating fully autonomously for any measured subset of AI R&D work." On more than 90% of research, Claude handles large chunks of the work under close human direction.

The distinction matters, and it cuts both ways. The generous read is that no researcher has been written out of the process. The harder read: the job moved from doing the work to directing and checking it, a different job with a different skill set, arrived at in half a year.

Here is what Anthropic put on the record.

What Anthropic measuredWhat it reported
Share of its AI R&D work Claude leads, February 2026Under 1%
Share of its AI R&D work Claude leads, August 202626%
Research where Claude handles large chunks under close human directionOver 90%
Agent decisions blocked by online monitors, August 2026About 0.002%, or 1 in 47,000
Transcripts flagged by offline monitors for human review1 to 2 per 1,000
High-priority flags escalated to a human reviewerAbout 50 per week

The oversight numbers are the more useful disclosure

Alongside the automation figures, Anthropic published how it watches its agents. Online monitors blocked about 0.002% of agent decisions across roughly a billion decisions analyzed in August, which works out to about 1 in 47,000. Offline monitors flag 1 to 2 transcripts per 1,000 for human review, and about 50 high-priority flags reach a human reviewer each week.

Most enterprises running agents today could not produce a single one of those figures about their own systems. That is the distance between deploying agents and actually governing them, stated in arithmetic rather than in principle.

Anthropic also proposed measures to carry the work forward. An R&D Automation Index would track how much of the research is AI-led. Agent Oversight Metrics would cover how much activity is monitored and how long review takes before something escalates. A Compute Allocation Measurement would show how much computing power goes to safety work against everything else. Dario Amodei has described the principle behind all of it in 5 words: "we must pace the frontier."

This is the half-life of skills with a number attached

I have been calling this the half-life of skills: the interval between learning something and that something going obsolete, and how that interval keeps shortening. It usually stays an argument, because the data to settle it does not exist. Companies do not measure how fast their own work is being absorbed. Anthropic just did, on itself, in public, on the one job nobody expected to go first.

Skills have expired this fast before. When sound arrived in film in 1927, performers with the wrong voice watched a craft they had spent careers building become unsellable in about 2 years. They were not bad at their work. The work stopped existing in the form they had mastered. The ones who came through it got to a microphone early, while there was still time to be a beginner.

Under 1% to 26% in 6 months is what the half-life of skills looks like when somebody finally puts a number on it.

Most companies are on this curve without a chart of it

Every company is somewhere on this curve. Anthropic happens to know where it stands on it. Your company sits on the same curve right now, in code, in legal review, in financial analysis, in customer support, and almost certainly has no idea what its own February-to-August number would be.

The measurement is not exotic, and I have started asking for it in my own advisory work. Pick one workflow that matters and write down what share of it a person still does end to end, against what share a person now directs and checks. Ask again in 90 days. That is a crude version of what Anthropic published, and it is enough to tell you whether your retraining budget is sized for the actual rate of change or for the rate you assumed 2 years ago.

This is what the agentic AI conversation has been missing while everyone argued about capability. Agents left the demo stage a while ago. The live question now is how quickly the work they absorb stops being work a person practices.

If you do not control a training budget, a smaller version of the question still belongs to you. What part of your week did you do entirely yourself a year ago that you now mostly review? That is your own number, nobody is going to calculate it for you, and it may be the most useful thing you can know about your next 5 years.

Multiply that across an economy and it stops being a career question. The rate at which skilled work gets absorbed decides whether the 51-year-old engineer gets a second act or a severance letter, and whether a household plans around one income or two. Anthropic's disclosure is one company's arithmetic. The pattern underneath it will reach a few million kitchen tables, and almost none of those households will get 6 months of warning with a chart attached.

Anthropic went from under 1% to 26% and knew it while it was happening. That is the whole advantage, and it costs a spreadsheet, not a research lab. So the question worth taking into your next leadership meeting is the one Anthropic answered about itself: what is our number, and when did we last check it? Amodei's line about pacing the frontier applies just as well to companies nowhere near the frontier. You cannot pace what you have never measured. And if you read this differently, I am easy to find.

Does a block rate of 1 in 47,000 mean the oversight is working?

On its own, that figure cannot tell you. A low blocking rate can mean the agents rarely attempt anything the monitors object to, or it can mean the monitors are set loosely. What makes the number readable is the second one published beside it: offline review flags 1 to 2 transcripts per 1,000 after the fact, which acts as a check on whether the live monitors are catching enough. Any company reporting a blocking rate without a separate after-the-fact review rate is reporting half a measurement.

Does this mean AI research jobs are going away at Anthropic?

Nothing in the disclosure says that, and no layoffs are part of this reporting. The measurement covers how tasks are distributed, not headcount. The change worth watching in figures like these is what the remaining work consists of, because supervising and directing draw on different skills than doing, and a company can hold headcount flat while the job under the title changes completely.

Should other companies expect to see a similar 26% figure in their own work?

Probably not, and the reason is worth knowing. Anthropic measured AI research performed by AI built for that research, inside a company with the engineering to instrument it, which is close to the fastest case available anywhere. Most sectors move slower, because their work still runs through regulated steps, physical processes, and records no machine can read. The number to compare against is your own rate of change over the same 6 months.

Here is what makes Alex a credible voice on this topic: Alex spent 20 years at Cisco, where he shaped a $1.1 billion innovation portfolio and watched automated systems absorb skilled work at scale, which is the curve Anthropic just published on itself.

To put a number on your own version of it, start a conversation →

← Back to Blog
Head portrait of Alex Goryachev
Alex Goryachev

WSJ-bestselling author · Former Managing Director of Innovation, Cisco · Advisor, CSU AI Working Group · LinkedIn Top AI Voice

Work with Alex

Bring this thinking to your organization

Alex works with executive teams at global enterprises on AI strategy, governance frameworks, and organizational readiness. Available for keynotes, C-suite workshops, and advisory engagements.