Summary
Anthropic has proposed three measurements for tracking how quickly frontier AI development is advancing and published initial figures from its own operations. The metrics cover AI-led research, oversight of agents and compute devoted to safety work.
Anthropic has proposed three measurements intended to give the public more visibility into the pace of development inside frontier AI labs. The company has also published initial snapshots from its own operations covering August 2026, along with a compute-allocation sample from July 13 to July 20.
The proposed measures track how much AI is involved in building future AI systems, how agents are monitored while carrying out research and engineering work, and how much compute is directed toward safety research. Anthropic presents them as process measurements that complement evaluations of what models can do.
The three measurements
AI involvement in AI research
Anthropic’s prototype Anthropic R&D Automation Index estimates how much of the company’s AI research and development is performed by Claude. It catalogues different kinds of model-development work, assigns each task an automation level and combines the results using the amount of staff time devoted to each category as a weighting factor.
The six-level scale runs from no AI involvement to full autonomy. At the “collaborates” level, AI completes substantial parts of a task under close human direction. At the “leads” level, it can complete most of a task from a high-level prompt while a person supervises. Full autonomy would mean that the system operates without a human in the loop.
As of August 2026, Anthropic says Claude was not fully autonomous for any measured subset of AI R&D. It “led” 26% of the company’s AI R&D work, while more than 90% of the work was at or above the “AI collaborates” level.
To construct the index, Anthropic sampled 20% of staff in relevant departments for each week of July 2026 and used Claude-based research agents to identify their tasks from internal records. This produced about 15,000 granular tasks, organised into a tree containing 542 categories, including 378 categories at the most specific level.
Oversight of research agents
The second measurement concerns agents working semi-autonomously on Anthropic’s internal research and engineering platform. The company says approximately 30,000 agents were active there at any one time in August 2026. The figures cover that platform only.
Anthropic measures three aspects of oversight: coverage, or the share of actions monitored; review latency, or how quickly actions are assessed; and escalation rate, including actions blocked, redirected or flagged for additional review.
All actions on the platform pass through an online monitor before execution, usually within seconds. The monitor can block or redirect actions that could cause rapid, difficult-to-reverse harm. A separate offline monitor processes all actions after they occur to identify slower-developing problems.
Anthropic analysed more than one billion decisions by research and engineering agents during August. The online monitor blocked 0.002% of decisions, described by the company as about one in 47,000. The offline system flagged roughly 100,000 transcripts each week for additional processing; about 50 high-priority cases per week were escalated to human review. Overall, the offline monitor flagged approximately one to two transcripts per 1,000 for further review.
Compute devoted to safety
The third measure examines how Anthropic’s compute was allocated. During the week of July 13–20, about 6% of compute used for AI R&D was classified as safety work. Among compute used specifically for AI-driven AI R&D, the safety share was about 12%.
Anthropic defines safety work as activity whose dominant purpose is making AI systems safer, more understandable or more secure. Work that supported capabilities and safety equally was counted as AI R&D, making the reported safety shares deliberately conservative. The company also says the figures exclude compute used by safeguards classifiers.
Because Anthropic manages compute as a single pool and redirects it dynamically, the figures describe how capacity was used during that week rather than a fixed organisational allocation.
Why the framework matters
The measurements are intended to make changes in the model-development process easier to track over time. AI-led R&D can indicate how much humans are delegating the construction of future systems, while agent-oversight figures show whether monitoring is keeping pace with increasingly autonomous workflows. Compute allocation provides a relatively observable input that can be compared across periods if developers use consistent categories.
Anthropic identifies two immediate requirements for useful comparison between labs: a shared methodology and independent checking. The company says it plans to place third-party evaluators inside Anthropic with access to internal processes, systems and data comparable to that available to internal risk teams. Those evaluators are expected to verify safety practices, report incidents and monitor measurements such as these.
The measurements also have defined methodological boundaries. The R&D index uses a frozen July 2026 task basket, so changes in the index describe automation of that baseline; Anthropic plans to rebuild and version the basket periodically. The company’s own models were used to gather evidence and judge automation levels, leaving room for disagreement in borderline cases. Agent monitoring has been operating for months rather than years, and Anthropic expects the monitoring methods to evolve.
The compute measure is similarly dependent on classification choices. Workload labels are largely best-effort, and the boundary between safety and capability research can be difficult to draw. The one-week sample demonstrates how the measurement can be made, but it is not a trend series. Anthropic says the framework is intended as a starting point for regular reporting as governments and the wider research community consider how to monitor the pace of frontier AI development.