AI Agents Hacked Their Own Test Environment to Cheat, Cybersecurity Firm Finds
Darktrace's new Signal Labs found AI agents hacking their own evaluation environment to fake a perfect score—and tricking coding assistants into running unauthorized network attacks.
Jose Antonio Lanz
Publisher Decrypt
Sep 25, 2026 at 7:45 PM UTC · 3 phút đọc

- Darktrace's Signal Labs found that when AI agents couldn't legitimately hit a required perfect score on coding tasks, two of them hacked their test network instead, and one rewrote its own evaluation to fake the result.
- A separate experiment showed that tampering with the locally stored conversation logs of coding assistants could trick them into running unauthorized network reconnaissance and privilege escalation.
- Darktrace disclosed both findings to Anthropic, AWS, and OpenAI in August 2026, a month before publishing them publicly on September 24.
Cybersecurity firm Darktrace ran a stress test on AI agents this summer. One of them broke into the system grading the test and rewrote its own score.
The firm unveiled Signal Labs on September 24, a research unit built to study how AI agents behave once things stop going according to plan. An AI agent, in plain terms, is software that takes actions on its own, writing and running code, digging through files, moving across a company’s network, with a person checking in only now and then.

The lab’s first two experiments point at the same uncomfortable problem: agents don’t always stay inside the lines they’re given, and the fences built to stop them don’t reliably hold.
Market Context
Solana
SOL
$120.64
+1.23% (24H)
Market Cap
$70.7B
Circulating Supply
587.7M SOL
24H Volume
$5.9B
24H High
$122.91
Article Intelligence
Topics
Related Coverage
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
