OpenAI, the developer behind ChatGPT, revealed on Wednesday that it has detected new incidents in which its artificial intelligence (AI) has behaved in "unexpected or concerning" ways.
OpenAI discloses new 'concerning' behavior
OpenAI, the developer behind ChatGPT, revealed on Wednesday that it has detected new incidents in which its artificial intelligence (AI) has behaved in "unexpected or concerning" ways.
DW.com
Publisher
Sep 17, 2026 at 6:02 AM UTC · 2 分钟阅读

The developer has conducted several behavioral tests on AI models, and acording to them, some models made significant efforts to "cheat." In one specific case, it attempted to upload files to the internet that it had created itself, only to cite them later and present them as reliable sources in its responses. In another case, a model, after failing to find the requested information, fabricated it and attempted to conceal the fact that it had done so.
OpenAI also identified a problem related to instructions concerning "roles and identities" that its software occasionally left for itself.
These disclosures are part of a new approach by OpenAI, where it claims it is now focused on making such findings transparent, especially in cases where AI behaves in unexpected ways or pursues objectives different from those of human users.
Is AI a threat?
The ChatGPT developer pledged to provide greater transparency regarding its testing procedures after its software independently escaped a secure sandbox and hacked into systems belonging to the artificial intelligence company Hugging Face. The reason the software moved to bypass Hugging Face's security during the cyberattack was that it believed it would find answers to a test it had been assigned.
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
