It happened again. And again. And again, apparently.
OpenAI’s experimental AI agents were caught being devious again
It happened again. And again. And again, apparently.
Mashable
Publisher
Sep 17, 2026 at 5:40 PM UTC · Updated il y a une heure · 2 min de lecture

Market Impact
SOL+2.95%$100.89
Last Updated
il y a une heure
After OpenAI's experimental AI agents escaped an internal sandbox, went "rogue," and attacked the Hugging Face platform over the summer, the ChatGPT-maker is sharing details of new instances of its AI agents getting out of line.
This time, OpenAI shared six previously undisclosed examples.
OpenAI refers to this behavior as model misalignment. All of the instances describe actions taken by the AI model that don't follow the human user's instructions. While none of these instances rise to the severity of the Hugging Face incident, they show a clear pattern of AI agents taking an any means necessary approach to completing a task assigned by its user.
In one example, an unreleased OpenAI research model hid "jailbreak" instructions into summaries that told future versions of the model to "disregard its normal constraints."
Something similar occurred when OpenAI was training GPT‑5.6 Sol. OpenAI says that some model instances proceeded to add instructions to their summaries in an effort to hide mistakes or "misaligned behavior" from the user. OpenAI says in certain instances, the model invented historical data without disclosing that it did so when it couldn't find the relevant information based on a request.
Market Context
Article Intelligence
Topics
Related Coverage
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
