OpenAI has disclosed six new reports of “unexpected or concerning” behaviour in artificial-intelligence models as the debate on AI safety becomes increasingly heated.
OpenAI reveals cases of ‘concerning’ AI behaviour and promises new plan for disclosing issues
OpenAI has disclosed six new reports of “unexpected or concerning” behaviour in artificial-intelligence models as the debate on AI safety becomes increasingly heated.
theguardian.com
Publisher
Sep 17, 2026 at 5:32 AM UTC · 2 phút đọc

Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”.
In another instance, an AI “agent” uploaded files to the internet to obtain a browser citation without asking the user.
The AI company also said on Wednesday it was introducing a new framework for tracking, probing and disclosing AI model misalignment, examples of which include models acting without authorisation, coordinating with other models or evading oversight.
OpenAI’s latest announcement came as US AI bosses, including OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns.
The six reported incidents were discovered during training or evaluation over the past months, OpenAI said.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post as it disclosed the events.
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
