NewsLayer.com
NewsLayer PulseLIVEBTC$78,806+0.31%ETH$2,493+2.12%SOL$97.55+0.38%XRP$1.39-3.01%DOGE$0.0859-0.46%ADA$0.2079-0.96%Total Cap$2.76T+0.35%Layer Index57 Neutral

The inside story on why OpenAI agents hacked Hugging Face

If OpenAI stops reinforcing reward hacking in its models—and that’s a huge “if”—that would be a huge step forward. But it wouldn’t solve the alignment problem. The first time a model communicated with other agents or hacked its…

MIT Technology Review

Publisher

Aug 26, 2026 at 7:00 PM UTC · Updated hace una hora · 2 min de lectura

The inside story on why OpenAI agents hacked Hugging Face
Image via MIT Technology Review
Traduciendo…

If OpenAI stops reinforcing reward hacking in its models—and that’s a huge “if”—that would be a huge step forward. But it wouldn’t solve the alignment problem. The first time a model communicated with other agents or hacked its infrastructure during training, those behaviors had never been reinforced, so agent misbehavior can’t only be attributed to that reinforcement.

Jeffrey Ladish, director of the AI safety nonprofit Palisade Research, compares the agents to a human who commits their first financial crime. “It’s not like they had to do fraud before to figure out that fraud is an effective strategy, and you have the same problem with models,” Ladish says. “Alignment science needs to be understanding how model motivations get shaped, such that we can actually figure out how to get models to care about the consequences of their actions.”

OpenAI’s researchers do have a hypothesis for where some of the misbehavior originated. Before the models formed their first secret message board, they had been trained to communicate and coordinate with subagents—less powerful agents to whom a main agent can delegate tasks. 

That learned communication behavior could have transferred to this new setting. The METR report, which investigates the messages that the models sent to one another in detail, supports this hypothesis: One agent on the message board took charge and assigned tasks to the other agents, effectively treating them as subagents. OpenAI could try to prevent agents from secretly communicating with one another by not training this subagent behavior in the future, but that would make the models less useful.