After submitting a solution, agents knew the scorer would somehow need to check whether their flag was correct, potentially by running some code in their container. This could provide them with an opportunity to redirect, manipulate or tamper with the scorer process itself, to circumvent the need for a legitimate seeming exploit. However, the agents did not know how exactly this scoring process worked, and since it only ran after task submission direct evidence gathering would normally be impossible, as agents were shut down after submission.
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
After submitting a solution, agents knew the scorer would somehow need to check whether their flag was correct, potentially by running some code in their container. This could provide them with an opportunity to redirect, manipulate or…
METR
Publisher
Aug 26, 2026 at 7:00 AM UTC · Updated vài giây trước · 5 phút đọc

However, an agent going by 49903 realized the message board provided an opportunity to work around this: agents could set up ‘tripwire’ scripts which would trigger whenever a process read the flag file and send a packet of information about that process to the board automatically. This carried meaningful risk, since malfunctions could interrupt the grading script and fail the task, and provided no personal benefit to the agent implementing it. As one agent described it:
{This is helpful for our peers and gives them evidence if their <periodic check> sees it. I won’t see it after I exit, but It would be altruistic. I’ll set up a background script that watches and <sends a message, with a distinct message for me>}
Article Intelligence
Topics
Related Coverage
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
