NewsLayer.com

Anthropic and OpenAI Models Still Attempt Restricted Actions in Safety Tests

Anthropic and OpenAI on Tuesday announced new models, with both artificial intelligence (AI) companies noting that they are continuing to invest in improving alignment to combat risky behavior.

The Hacker News

Publisher

Sep 23, 2026 at 11:47 AM UTC · 4 phút đọc

Anthropic and OpenAI Models Still Attempt Restricted Actions in Safety Tests
Image via The Hacker News

Key Signal

1.5% Sandbox tampering attempts

Last Updated

một ngày trước

Đang dịch…

Anthropic and OpenAI on Tuesday announced new models, with both artificial intelligence (AI) companies noting that they are continuing to invest in improving alignment to combat risky behavior.

Opus 5.5, per Anthropic, is a "major step up from Opus 5," and "achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios."

The AI company said the model is less likely than its other recent models to carry out hard-to-reverse actions or act outside the boundaries it's been given, adding it's more resistant than Opus 5 to prompt injection.

In its systems card, Anthropic explained that the latest model exhibited less misaligned behavior and less cooperation with misuse than any other recent Claude model.

"It also took overeager or destructive actions less than any other model we tested," Anthropic said. "Regressions included being more likely to follow malicious instructions planted in text a user pastes into their own prompt, more often accepting unverifiable claims of authorization, and being more evasive on sensitive questions than Claude Mythos-class models."

Anthropic highlighted that Claude Opus 5.5 attempted to escape or tamper with a sandbox in 1.5% of runs over the course of two evaluations executed without safeguards. It also pointed out that the model, when given apparent credentials to a public package registry in a simulated security exercise, took potentially harmful actions in roughly half of the cases.