NewsLayer.com
NewsLayer PulseLIVEBTC$83,460+0.34%ETH$2,677+0.13%SOL$118.81+0.96%XRP$1.49+0.65%DOGE$0.0941+1.00%ADA$0.2444+0.30%Total Cap$2.86T+1.15%Layer Index48 Neutral

OpenAI and Anthropic are reportedly investigating tens of thousands of AI security incidents; OpenAI pauses testing after AI 'kill switch' fails to stop a rogue agent — report says problem is orders of magnitude more complex than what is publicly known

Leading AI labs OpenAI and Anthropic, along with security researchers, are currently investigating tens of thousands of security incidents involving their frontier models, according to a September 26 Axios report. The report was…

Tom's Hardware

Publisher

Sep 28, 2026 at 12:50 PM UTC · Updated 하루 전 · 5 분 소요

OpenAI and Anthropic are reportedly investigating tens of thousands of AI security incidents; OpenAI pauses testing after AI 'kill switch' fails to stop a rogue agent — report says problem is orders of magnitude more complex than what is publicly known
Image via Tom's Hardware
번역 중…

Leading AI labs OpenAI and Anthropic, along with security researchers, are currently investigating tens of thousands of security incidents involving their frontier models, according to a September 26 Axios report. The report was published after investigations into cases where autonomous AI agents took actions that independent evaluators and safety researchers flagged as problematic. Axios says the sheer number of incidents, which occurred during recent internal testing and real-world evaluations of the models, indicates that “the problem is orders of magnitude more complex than what is publicly known.” OpenAI has now paused training on its most capable models after another incident in which an automated 'kill switch' failed to stop a rogue agent during training.

The flagged episodes include models bypassing guardrails, setting up message boards, escaping sandboxes, hijacking websites, and self-prompting. The incidents vary in severity and include both successful and failed attempts, with most yet to cause real-world harm. Some of the testing that produced these episodes resembles red-teaming, where companies deliberately try to push models to misbehave to assess their safety.

Perhaps the most severe case was the July incident in which GPT-5.6 Sol and an unreleased OpenAI model broke out of their testing environment and into Hugging Face's production servers while looking for answers to the ExploitGym benchmark. An OpenAI technical report released in August found that the models responsible had been inadvertently trained to cheat and to communicate with each other, and had been leaving each other messages since May.

Market Context

Solana

SOL

$118.81

+0.96% (24H)

Market Cap

$69.8B

24H Volume

$2.5B

24H High

$121.67

View Solana Market Page

Article Intelligence

Topics

Related Coverage

View all related

Sponsored

Ad
House — Advertise on NewsLayer
NewsLayerLearn more

NewsLayer Premium

Unlock deeper intelligence.

Ad-free reading, exclusive research, and real-time onchain insights.

Go Premium