Anthropic is proposing three new metrics that it says could help the artificial intelligence industry measure and monitor the pace of frontier AI development, days after CEO Dario Amodei called for a coordinated slowdown in the advancement of sophisticated models.
The company published the measurements Thursday, covering AI-assisted research and development, oversight of AI agents and the allocation of computing resources. Anthropic also released details of its methodologies, encouraging other AI companies to adopt similar measurements and make their results public.
The initiative builds on a three-step slowdown plan Amodei published Saturday. While that proposal called for greater coordination around the pace of frontier AI development, it provided limited detail on how such a system could be implemented in practice.
Register for the next Tekedia Mini-MBA.
Register for Tekedia AI in Business Masterclass.
Join Tekedia Capital Syndicate and co-invest in great global startups.
Anthropic’s latest proposal is an attempt to put measurable indicators around that discussion.
“As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows,” Anthropic said in its blog post. “This means better measuring the development of AI, reporting on it publicly, and giving society an opportunity to decide how to use this information.”
The timing puts Anthropic at the center of an active debate over whether frontier AI development should continue at the current pace or be accompanied by stronger monitoring and safety measures.
Amodei’s call for a slowdown received support from several technology executives, including OpenAI CEO Sam Altman, SpaceX CEO Elon Musk and Google DeepMind Chair Demis Hassabis. His proposal followed warnings from AI researchers about the potential risks associated with more capable models.
Amodei has said his approach is intended to slow the rate at which model capabilities improve without sacrificing commercial competitiveness or the United States’ position in AI.
Measuring How Much AI Is Involved In AI Development
Anthropic’s first metric examines the extent to which its own Claude models are capable of autonomously carrying out research and development work. The company said that, within the subset of research and development activities it measured, Claude models were “not operating fully autonomously.”
The measurement is intended to address a question that could become more important as AI systems are used to develop subsequent generations of AI: how much of the work required to advance AI is itself being performed by AI?
As models become more capable at coding, experimentation, analysis, and other technical tasks, the boundary between AI-assisted development and AI-driven development could become difficult to define without standardized measurements.
Anthropic’s second metric focuses on oversight of AI agents.
The company built a system designed to monitor and intervene in actions taken by AI agents and found that approximately 30,000 agents were conducting research and engineering work across its most-used internal platform at any given time.
The figure illustrates the scale at which agentic AI is already being incorporated into internal technical workflows. Rather than measuring only the capabilities of individual models, the metric looks at the number of AI agents operating within an organization’s development environment and the systems used to supervise them.
AI systems are given greater autonomy and access to software tools, internal information, and computing resources, making the idea relevant.
Anthropic Tracks AI Safety Spending
The third metric examines how Anthropic allocates its computing resources. The company conducted a snapshot of its total compute usage between July 13 and July 20. It found that approximately 6% of the compute used for AI research and development was allocated to safety work.
When measured specifically against compute allocated to “AI-driven” research and development, Anthropic said about 12% was directed toward safety.
The figures provide a quantitative measure of the resources Anthropic is allocating to safety relative to model development. They also establish a baseline that other AI companies could potentially replicate, allowing comparisons across laboratories if similar definitions and measurement methodologies are adopted.
Anthropic stressed that the three metrics are not intended to replace existing capability evaluations. Instead, the company said the measurements should complement evaluations that show “what models can do” by providing information about how models are developed and how much AI is involved in that development.
Together, Anthropic said, the measurements could provide organizations outside AI laboratories with a starting point for assessing the pace of frontier AI development.
The proposal comes at a time when the industry is focused not only on the capabilities of AI models but also on the speed at which those capabilities are improving. As laboratories compete to build more powerful systems, policymakers and the public have limited visibility into the amount of computing, automation, and human oversight involved in that process.
Anthropic’s argument is that greater disclosure could narrow that information gap.
“We hope to model that transparency by releasing these measurements, and we’ll continue to do so,” the company said.
However, it is believed that the effectiveness of the framework will depend partly on whether other frontier AI developers adopt comparable measurements. Without common definitions and reporting standards, figures such as the share of compute devoted to safety or the number of active AI agents could be difficult to compare across companies.



