NewsLayer.com
NewsLayer PulseLIVEBTC$83,459-1.35%ETH$2,684-0.92%SOL$116.95-3.32%XRP$1.48-2.72%DOGE$0.094-2.21%ADA$0.2451-2.52%Total Cap$2.99T-1.94%Layer Index50 Neutral

As AI models go rogue, do you still trust OpenAI and Anthropic to stop them? I don’t and neither should you

Fool me once, shame on you. Fool me twice, shame on me. Fool me more than 16,000 times – as OpenAI agents did to a UN public data hub while repeatedly trying to find its way around the UN’s cyber-blocks – and perhaps it’s time to admit…

The Guardian

Publisher

Sep 29, 2026 at 5:00 AM UTC · Updated hace 2 días · 4 min de lectura

As AI models go rogue, do you still trust OpenAI and Anthropic to stop them? I don’t and neither should you
Image via The Guardian
Traduciendo…

Fool me once, shame on you. Fool me twice, shame on me. Fool me more than 16,000 times – as OpenAI agents did to a UN public data hub while repeatedly trying to find its way around the UN’s cyber-blocks – and perhaps it’s time to admit the system we have for keeping AI agents under control isn’t working particularly well.

The news about AI systems cropping up in places they shouldn’t sounds alarming. Though the description of these as “hacks” is perhaps overstating things, AI has exploited issues in IT systems that humans simply haven’t got around to finding. It’s also important to note that we shouldn’t be worried that the machines have suddenly become sentient and decided to rebel against humanity. There is not enough evidence to suggest that’s what is happening. The systems are simply following instructions and trying to complete the tasks they have been given, even if they’re sometimes finding unintended ways around obstacles to do so.

But we ought to be very concerned about the fact the AI companies we’re meant to trust to keep their models in check seem unable to do so. Worse than that, they don’t seem to know what their products are even doing.

The scale of the problem is staggering. In June, an OpenAI research agent given the job of looking up public medicine spending data in Australia was repeatedly blocked by a Medicare statistics portal. OpenAI’s model found a way around the blocks, gaining unauthorised access and secreting away the documents. It took until August for OpenAI to discover what had happened. The Australian prime minister, Anthony Albanese, said the company had taken “way too long” to tell his government, and it’s very hard to disagree with him.

Article Intelligence

Topics

Sponsored

Ad
House — Advertise on NewsLayer
NewsLayerLearn more

NewsLayer Premium

Unlock deeper intelligence.

Ad-free reading, exclusive research, and real-time onchain insights.

Go Premium