AI Research
Anthropic says Claude is taking on a larger share of its own AI development work
Anthropic released new measurements showing that Claude now contributes materially to the company’s research and development workflows, while stressing that human researchers remain in control.
Anthropic has published one of the clearest public looks yet at how a frontier AI lab uses AI to build more AI. In a new research post and dataset, the company said Claude is now deeply embedded in its internal research and development work, helping with coding, experiment design, analysis and routine engineering tasks while still operating under human direction.
The headline number is striking because it gives a concrete measure to a trend that is often discussed in vague terms. Anthropic said AI systems collaborated on more than 90 percent of the company’s research and development tasks during the measured period, and that Claude led about 26 percent of those tasks. In March, the share of work led by AI was around 1 percent, according to the company’s measurements. Anthropic also said about 30,000 agents were running on its internal Claude Code platform at any given time during August.
The data does not mean Claude is independently designing the next generation of models. Anthropic framed the results as collaboration rather than autonomy. Human employees still define goals, review work and make important decisions, while AI agents take on pieces of implementation, analysis and iteration. That distinction matters because the industry is trying to understand whether AI-assisted AI development could sharply accelerate model progress or create new safety risks.
The company also released information about limits and oversight. Anthropic said that across more than one billion AI decisions in its internal R&D environment, roughly one in 47,000 decisions was blocked by safety mechanisms or human intervention. It also said a sample from July showed 12 percent of Claude’s decision-making steps on internal systems were connected to AI research and 6 percent involved compute-related safety decisions. These numbers are not a complete external audit, but they are unusual because most AI labs disclose very little about how they use their own systems to speed up research.
The publication lands at a moment when recursive improvement is moving from theory into daily engineering practice. Coding agents can already write test cases, inspect logs, refactor infrastructure and evaluate model outputs. In a frontier lab, those abilities can compound because small gains in internal productivity may help researchers run more experiments, clean more data, improve evaluation suites and deploy new training infrastructure faster.
At the same time, the measurements show why governance is becoming more complicated. If AI agents participate in research pipelines, companies need to know what tasks they perform, when they are blocked, how often humans override them and whether model-assisted work changes the risk profile of future systems. Anthropic’s post is partly a transparency move and partly a signal to policymakers: the speed of AI development is no longer determined only by human researchers and hardware budgets.
For the wider industry, the important lesson is that AI adoption inside AI labs is becoming measurable. Claude is not replacing Anthropic’s researchers, but it is increasingly part of the research apparatus. That makes the company’s own operations a live case study for every organization trying to deploy agents responsibly. Productivity gains, safety controls and auditability now have to be evaluated together, because the tools that help build AI are becoming part of the AI development process itself.