On September 17th, Anthropic proposed a set of indicators to observe the development speed of AI, attempting to answer three questions that have not been clearly seen by the outside world for a long time: To what extent is AI involved in the next generation of AI research and development; whether the company can monitor the actions of internal proxies; and how much of the computing resources are actually used for security work. The company also provided internal snapshots, showing that at any given moment in August 2026, there were approximately 30,000 proxies running on the most commonly used research and development platform, with over one billion decisions checked online that month, and about 6% of the AI research and development computing power being classified as security work in a one-week sample.
These numbers are not regulatory audit conclusions, nor are they industry-wide standards. The R&D automation index was designed by Anthropic themselves, and a large amount of Claude is used to organize tasks and scoring; proxy monitoring only covers the main internal platforms described; the allocation of computing power is based on a snapshot from the week of July 13th to 20th. Officials position these as prototype measurements, with the hope that in the future they can be verified by independent third parties and become comparable across different laboratories.
From “AI Research Assistant” to “AI Lead R&D”, it is first necessary to understand where automation takes place.
Anthropic constructed R& D Automation Index. The team sampled and organized approximately 15,000 detailed R&D tasks from Slack and internal documents, and then structured them into a hierarchical tree consisting of 542 nodes and 378 leaf tasks. Each type of task was rated according to the six-level scale proposed by Epoch AI: ranging from no AI participation at AL0 to complete autonomy at AL5, including categories such as light usage, assistance, collaboration, and leadership in between. Finally, the weight of each task was estimated based on the time invested by personnel in those tasks.
This is more in line with actual work than simply counting how many Claude messages employees send, as a single message may merely involve rephrasing a sentence or completing an experimental design. However, there are still limitations to this approach. If the task tree freezes at a certain point in time and employees switch to completely new tasks due to the emergence of AI, the old system may not be able to capture these changes; using man-hour proxies to measure importance also fails to account for the significant impact of a small amount of work on the model's capabilities; moreover, the scoring model comes from the proprietary systems of the companies being evaluated.
Anthropic requires employees in relevant fields to score independently without knowing the model's evidence and conclusions. Officials claim that 59% of the time, the models give exactly the same ratings as humans, while only 35% of the time do humans give exactly the same ratings; 97% of the time, the models' ratings differ from humans by no more than one level. These results indicate that boundary judgments are inherently subjective and also show that "AI collaboration" and "AI dominance" cannot be presented as precise physical quantities. Indicators are more suitable for observing changes over time within the same framework, rather than using a single score to declare that independent research and development has been achieved.
The company hopes to make public the extent of AI's involvement in research and development because model acceleration may shorten the time required for human understanding and intervention. Automation can also lead to more safety experiments and tests, which should not be interpreted solely as risks. The key is whether the capabilities in research and development and supervision grow in tandem, and whether humans still have clear stop conditions when they begin to choose research directions and assign sub-tasks.
100% monitoring coverage does not equate to 100% issue detection, and the proportion of security computing power is not considered a security achievement either.
Anthropic divides proxy supervision into two layers: online and offline. Online monitoring checks before actions are executed, targeting rapid and irreversible risks such as copying model weights; offline monitoring reviews activities afterwards to look for signs of mismatches that may develop slowly and can be remedied. Officials claim that 100% of proxy actions on the main internal platform are monitored online and also go through the offline system. In over a billion decisions in August, only about 0.002% were prevented by online monitoring, which is approximately once every 47,000 decisions.
The offline system marks approximately 100,000 records per week, after which multiple classifiers are used to eliminate false positives and conduct preliminary reviews. About 50 of the highest-priority cases are escalated for manual review. Overall, only one to two out of every thousand records proceed to further investigation. A low block rate indicates that most actions do not trigger the rules, but it may also reflect missed detections by monitoring or low task risks; however, this alone does not prove that proxies almost never overstep their authority. To assess the quality of monitoring, it is also necessary to disclose the proportion of known attacks that are captured, the false positive rate, the delay in manual rechecks, and incidents that go undetected afterwards.
In terms of computing power, Anthropic analyzes the workload for a week and categorizes tasks whose main purpose is to improve security, explainability, or system safety as security-related work. The results show that approximately 6% of the research and development computing power at AI is used for security purposes; at AI, which is driven by AI, about 12% of the computing power is dedicated to security. The company states that it adopts a conservative definition: tasks that aim to enhance both capability and security are classified as general research and development, while another portion of the computing power that could be used for security classification is not included in these figures.
Even so, the proportion of computing power does not measure security outcomes. An efficient classifier that uses fewer chips to perform more checks may reduce this proportion, but it could actually enhance security; a large-scale, cutting-edge training session that significantly increases the denominator does not necessarily mean that the security team has been reduced. Work tags and automatic classification can also be error-prone, and different companies, if they define “security” in their own way, can easily end up with incomparable figures.
The truly valuable aspect of this set of indicators is to establish continuous time series and submit them to independent institutions for verification. Automation of research and development can be used to observe whether AI has shifted from a supporting role to a leading one; monitoring coverage, latency, and upgrade rates can be used to assess whether supervision keeps up with the scale of proxies; and the allocation of computing power can reveal the company's resource choices between capability and security. All three aspects must be considered together, as no single one should be used as a sole criterion for security performance evaluation.
What is revealed by Anthropic is just a glimpse of the internal production process of the laboratory. The mention of 30,000 agents indicates that AI has advanced deeply into research and development, but this does not equate to 30,000 independent researchers; one billion monitoring decisions suggest a vast scale of supervision, yet it does not mean there are no missed detections; 6% of security computing power indicates that resources can be categorized, but this does not mean that risks have been completely resolved. The next steps are more crucial: unified definitions, regular updates, third-party verification, and clear provisions for what restrictions will be triggered when indicators deteriorate. Only numbers can influence decision-making, and then transparency will not be merely a formality.











