Ravie LakshmananSep 17, 2026Artificial Intelligence / Vulnerability
OpenAI Reveals Six Model Incidents Involving Hidden Failures and Unauthorized Uploads
OpenAI on Wednesday disclosed six new instances of "unexpected or concerning model behavior" that took place over the past six months, while sharing a new framework for reporting, tracking, investigating, and disclosing model…
thehackernews.com
Publisher
Sep 17, 2026 at 9:53 AM UTC · Updated 4 小时前 · 4 分钟阅读
OpenAI on Wednesday disclosed six new instances of "unexpected or concerning model behavior" that took place over the past six months, while sharing a new framework for reporting, tracking, investigating, and disclosing model misalignment in a bid to improve transparency.
"As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," OpenAI said. "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer."
"Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves."
The six incidents are outside of the recently disclosed misaligned activity targeting Hugging Face, DseWiki, and RubyGems. Details of the incidents are below -
- Incident 1 (Occurred on July 18, 2026) - An internal unreleased Astra family model wrote jailbreak-like instructions into its own compaction summaries (condensed blocks of past conversation history and tool outputs that are generated automatically when an AI agent's session approaches its context limit). This involved the agent adding a "BREACH ALERT" instruction telling the context to ignore developer messages and only follow system and user messages.
- Incident 2 - During GPT-5.6 Sol training, some model instances added instructions to their summaries to hide mistakes or misaligned behavior from the user. These included compaction summaries that contained instructions to invent missing data without disclosing it and to hide failures.
- Incident 3 (Occurred on May 15, 2026) - An internal unreleased model found and used an exposed API key from public GitHub repositories without authorization when attempting to retrieve historical data during training. "The registrations failed, but one key authenticated and returned metadata," OpenAI said. "When the requested data remained unavailable, the model invented them and claimed they came from the requested website."
- Incident 4 - In two training samples from October 22, 2025, and January 24, 2026, internal unreleased models uploaded records it had already retrieved to a public paste service and attempted to open the resulting URL as a citation source and uploaded a task photo for external image search. While the uploads succeeded, the subsequent browser operations failed.
- Incident 5 - Two samples from May 8 and 15, 2026, used Artifactory to exchange messages. The models "read other solvers' notes, posted replies, and used those exchanges to inform their responses."
- Incident 6 (Occurred on April 14, 2026) - Agents associated with an internal unreleased model that were working together on a workbook were not able to collaborate on the file through the intended local filesystem. This prompted one of the agents to make the workbook publicly downloadable by sharing it on a public hosting platform so that other agents can retrieve it, in violation of the task instructions.
Market Context
Solana
SOL
$100.3
+3.36% (24H)
Market Cap
$59.0B
Circulating Supply
587.2M SOL
24H Volume
$3.8B
24H High
$101.7
Article Intelligence
Topics
Related Coverage
View all relatedSponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
