Gemini 4 Is Here, and Google’s Flagship Tops All Other AI Models on Cybersecurity
Gemini 4 Argon tops 12 of 18 benchmarks in Google's own table, writes a million tokens per reply and resists hijacking best. Cyber defenders get it first, with the guardrails off.
Jose Antonio Lanz
Publisher Decrypt
Sep 30, 2026 at 11:20 PM UTC · 4 min de leitura

Key Signal
0.7% Prompt injection attack success
Market Impact
SOL-0.87%$117.75
Last Updated
há 15 horas
- Google unveiled Gemini 4 Argon on Wednesday, scoring 77.9% on DeepSWE v1.1 and leading 12 of 18 benchmarks in its own comparison table.
- It posted a 0.7% attack success rate on Gray Swan's prompt injection test, ahead of Claude Opus 5.5 and Claude Fable 5.1, which both scored 1.0%.
- Argon goes first to vetted cyber defenders through the Fairwind Program, without cyber guardrails, before reaching paid API customers and Google AI Ultra subscribers.
Gemini 4 is finally here, one week after the release of Claude Opus 5.5 and one day after GPT 6.1 Sol, proving American labs are very much committed to slowing down AI development. Please excuse our sarcasm.
Google unveiled Gemini 4 Argon on Wednesday, calling it its frontier model, meaning its most capable, for coding, office work and cyber defense.

On DeepSWE v1.1, a test of whether an AI can finish long, messy, real-world software engineering jobs, scored as a percentage, Argon hit 77.9%. Claude Opus 5.5 got 74.2%, GPT-6 Astra 74.1% and Claude Fable 5.1 67.4%.
For scale, Gemini 3.6 Flash managed 49% on the same test in July. Argon can also write up to 1 million tokens in one reply, up from 64,000. A token is a chunk of text, roughly three-quarters of a word, so that is about 750,000 words versus about 48,000.
Market Context
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
