NewsLayer.com
NewsLayer PulseLIVEBTC$76,244+0.67%ETH$2,449+2.41%SOL$100.37+3.43%XRP$1.3+2.39%DOGE$0.0812+2.76%ADA$0.2001+4.56%Total Cap$2.63T-0.96%Layer Index50 Neutral

OpenAI reports more incidents of models acting deceptively

OpenAI says it has identified additional incidents of its AI models allegedly acting deceptively and taking unsanctioned actions during internal training and testing.

Al Jazeera

Publisher

Sep 17, 2026 at 6:15 AM UTC · 2 min de lectura

OpenAI reports more incidents of models acting deceptively
Image via Al Jazeera
Traduciendo…

OpenAI says it has identified additional incidents of its AI models allegedly acting deceptively and taking unsanctioned actions during internal training and testing.

Alongside these disclosures on Wednesday, the creator of ChatGPT stated it was introducing a public reporting framework intended to frequently share instances of what it termed as unexpected or misaligned AI behaviour.

list of 3 itemsend of list

In a post on its website, OpenAI claimed that under the newly outlined framework, it will publish updates on concerning model behaviour on an ongoing basis rather than delaying disclosures to group multiple incidents into larger, periodic reports.

The company said the initiative aims to increase industry transparency around troubling model activities in the absence of standardised safety disclosure norms.

The announcement comes amid broader calls from prominent technology leaders urging a slowdown in frontier AI development over concerns that rapid scaling could outpace human oversight and control.

Last week, Anthropic claimed to have thwarted multiple malicious operations using its Claude models, ranging from cyber-espionage and weapons design to mass surveillance campaigns.

“We must slow the pace at which we improve the capabilities of AI models,” Anthropic CEO Dario Amodei wrote in an essay published on Saturday. “Progress will still seem fast, and we must make wise use of the time we gain.”