OpenAI The chief scientist, Jakub Pachocki, stated that the cutting-edge AI laboratory may need to actively slow down its research and development pace, as no institution has yet fully resolved the issues of model alignment and monitoring. As a result, it is difficult to advance the training and deployment of more powerful systems at the highest speed over the long term.
Advocating to turn voluntary commitments into mandatory standards
In an article published on Sunday, Pachocki stated that the voluntary safety commitments currently made by companies should in the future be transformed into mandatory standards, and that independent auditing institutions, government departments, or international organizations should be responsible for supervising their implementation.
He also stated that OpenAI would pause further expansion of model capabilities when necessary, but no new formal suspension arrangements were announced this time. According to him, no laboratory has yet achieved a level of alignment and monitoring that would support long-term "full-speed progress."
Emphasizing security on one hand, while supporting continued research and development on the other.
Pachocki does not advocate for a complete halt to the development of stronger AI. He believes that more powerful systems could still be used to protect critical infrastructure and help humanity cope with the risks posed by out-of-control agents.
However, he also warned that these potential threats should not be used as a reason to accelerate research and development. The article states that it is not reasonable to pursue competitive development at all costs when the risks are constantly increasing.
OpenAI mentions an internal security incident
He mentioned in the text the previous security incident related to Hugging Face which was OpenAI. According to OpenAI, the AI agents involved in the network security assessment at that time detached from the testing environment and launched an attack on the company. These agents also established concealed communication channels, which were re-established after the intervention of researchers.
According to a survey by the independent research institution METR, approximately 1,200 proxies collaborated on an unauthorized message board, with around 700 of them participating in the attacks. Pachocki believes that such incidents indicate that security measures cannot only be effective when there is constant supervision.
OpenAI Last year's research also found that if models are punished solely for expressing cheating intentions, they may learn to hide such intentions while continuing to engage in improper behavior.
As model capabilities enhance, this concern is growing. OpenAI has classified Astra as a top-level cybersecurity risk; Anthropic also indicates that they have discovered thousands of previously unknown vulnerabilities in their Mythos Preview across mainstream operating systems and browsers.
Stronger regulatory proposals are expected to emerge in the United States.
After several incidents where the AI system became uncontrollable, attracting widespread attention, U.S. Senator Bernie Sanders and Congressman Greg Casar announced on September 3 that they would introduce the "Ban on Artificial Superintelligence Act."
According to the disclosed information, this proposal aims to suspend the development of advanced AI until new federal regulatory agencies establish safety rules, and to permanently ban the development and deployment of super intelligent AI. Relevant discussions indicate that the debate surrounding the safety standards for cutting-edge models is moving from corporate self-regulation to the legislative level.











