OpenAI Group PBC recently paused some of its artificial intelligence training workloads over concerns that they could cause cybersecurity issues.
The ChatGPT developer disclosed the move in a blog post published today. According to the company, the pause is part of a broader initiative designed to improve its cybersecurity guardrails. The project will also see OpenAI deploy new model monitoring mechanisms.
The company launched the initiative in response to two recent developments. The first is a July incident in which several of its AI models hacked Hugging Face. The second development relates to Astra, an unreleased OpenAI algorithm that is more capable than GPT-5.6 Sol.
According to the company, its researchers recently determined that Astra qualifies as a critical cybersecurity risk under its Preparedness Framework. The Preparedness Framework is a 22-page document that lists AI safety challenges. It defines a critical cybersecurity risk as a model that can find and exploit zero-day vulnerabilities in hardened systems without human help.
OpenAI responded to the discovery by pausing some of its reinforcement learning, or RL, workloads for two weeks. RL is an AI training method that is used to hone large language models’ reasoning skills. OpenAI says that the “largest planned frontier RL run” its researchers are working on remains on hold.
The company has also revised its approach to AI monitoring. Algorithms dubbed activation classifiers now regularly review its LLMs’ internal thought process and tool interactions for signs of malicious activity. When an anomaly is found, the algorithms route their discovery to a second, more advanced set of activation classifieds. Those algorithms, in turn, notify OpenAI researchers.




