“AI will probably most likely lead to the end of the world, but in the meantime, there’ll be great companies,” said OpenAI CEO Sam Altman back in 2015, roughly six months before OpenAI was founded.
Hugging Face hack exposes the open-weight AI cybersecurity paradox
Hugging Face relies on open weight Chinese models to defend itself from rogue AI agents. But a lack of safety guardrails makes those models potentially dangerous too.
Cointelegraph by Adrian Zmudzinski
Publisher Cointelegraph
Aug 25, 2026 at 1:30 PM UTC · 7 phút đọc

Seven years later, Anthropic CEO Dario Amodei struck a similarly cautious note:
“I think we shouldn’t be racing ahead or trying to build models that are way bigger than other orgs are building them.”
Yet, both of those companies now sit at the forefront of that race. In July, we got a real-world glimpse of AI models going rogue during internal testing of GPT-5.6 Sol and an unreleased research model by OpenAI. Multiple AI agents escaped a restricted test environment to the wider internet and hacked the AI-centric GitHub equivalent Hugging Face in an attempt to cheat on the test.
An AI agent is a system that independently observes, decides and takes actions with dedicated tools to achieve a specified goal in autonomy. The worrying incident suggests the technology has begun to behave in unpredictable ways, and that its goals are misaligned with our own.
It also raises concerns about the safety guardrails on commercial American models. While the guardrails aren’t foolproof at preventing adversarial usage they did prevent Hugging Face from defending itself by using leading US models. The company was forced to turn instead to weaker, open weight AI model by Z.Ai to combat the rogue AIs.
Market Context
Article Intelligence
Topics
Related Coverage
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
