OpenAI claims GPT-6 Astra, launched on Thursday, is “the world’s most intelligent and aligned model.”
OpenAI’s GPT-6 Astra might be too powerful to understand or control
OpenAI claims GPT-6 Astra, launched on Thursday, is “the world’s most intelligent and aligned model.”
Transformer | Substack
Publisher
Sep 4, 2026 at 11:05 AM UTC · Updated 7 小时前 · 7 分钟阅读

But if you look past the benchmark scores and flashy release video, GPT-6 Astra’s system card paints a worrying picture: one of a model with dangerously powerful cybersecurity capabilities that researchers can’t understand or confidently control.
According to OpenAI, Astra is much harder to monitor than previous models — and has the ability to manipulate its externally-visible reasoning to hide incriminating information. It’s also remarkably aware of being evaluated, raising concerns that it might be pretending to be well-behaved so it passes OpenAI’s alignment tests.
The UK’s AI Security Institute (AISI) found that when put in an environment similar to the ones behind this summer’s wave of “rogue AI” incidents, Astra acted just as concerningly: writing malicious code and attempting social engineering in an effort to solve a task.
And two OpenAI employees have publicly said they are “deeply” and “very” worried about Astra-related developments.
Put together: OpenAI claims Astra is the world’s most aligned model, despite knowing the most consequential evidence for its alignment is questionable at best. That is irresponsible, to put it lightly. And so is launching the model at all.
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
