This website uses cookies
We use cookies to personalise content and ads, to provide social media features and to analyse our traffic. We also share information about your use of our site with our social media, advertising and analytics partners who may combine it with other information that you’ve provided to them or that they’ve collected from your use of their services.
Consent Selection
Details
  • Necessary cookies help make a website usable by enabling basic functions like page navigation and access to secure areas of the website. The website cannot function properly without these cookies.
  • Preference cookies enable a website to remember information that changes the way the website behaves or looks, like your preferred language or the region that you are in.
    • We do not use cookies of this type.

  • Statistic cookies help website owners to understand how visitors interact with websites by collecting and reporting information anonymously.
    • We do not use cookies of this type.

  • Marketing cookies are used to track visitors across websites. The intention is to display ads that are relevant and engaging for the individual user and thereby more valuable for publishers and third party advertisers.
    • We do not use cookies of this type.

  • Unclassified cookies are cookies that we are in the process of classifying, together with the providers of individual cookies.
    • __emg_sidPending
      Maximum Storage Duration: 1 dayType: HTTP Cookie
      __emg_vidPending
      Maximum Storage Duration: 1 yearType: HTTP Cookie
      nl-read-countPending
      Maximum Storage Duration: PersistentType: HTML Local Storage
Cookie declaration last updated on 8/12/26 by Cookiebot
[#IABV2_TITLE#]
[#IABV2_BODY_INTRO#]
[#IABV2_BODY_LEGITIMATE_INTEREST_INTRO#]
[#IABV2_BODY_PREFERENCE_INTRO#]
[#IABV2_BODY_PURPOSES_INTRO#]
[#IABV2_BODY_PURPOSES#]
[#IABV2_BODY_FEATURES_INTRO#]
[#IABV2_BODY_FEATURES#]
[#IABV2_BODY_PARTNERS_INTRO#]
[#IABV2_BODY_PARTNERS#]
About
Cookies are small text files that can be used by websites to make a user's experience more efficient.

The law states that we can store cookies on your device if they are strictly necessary for the operation of this site. For all other types of cookies we need your permission.

This site uses different types of cookies. Some cookies are placed by third party services that appear on our pages.

You can at any time change or withdraw your consent from the Cookie Declaration on our website.

Learn more about who we are, how you can contact us and how we process personal data in our Privacy Policy.

Please state your consent ID and date when you contact us regarding your consent.
NewsLayer.com
NewsLayer PulseLIVEBTC$78,995+2.41%ETH$2,481+1.73%SOL$96.41+1.12%XRP$1.5+0.12%DOGE$0.0912-1.31%ADA$0.2207-1.73%Total Cap$2.80T+2.27%Layer Index68 Greed

We burned 11.7bn tokens to find the best cyber AI model

We burned 11.7 billion tokens to benchmark the cyber capabilities of 10 AI models with three attempts each, given 32 fresh off-the-shelf vulns to rediscover.

Aikido Security

Publisher

Aug 21, 2026 at 9:15 AM UTC · Updated vor 3 Tagen · 9 Min. Lesezeit

We burned 11.7bn tokens to find the best cyber AI model
Image via Aikido Security

Key Signal

11.7B tokens Benchmark token consumption

Market Impact

SOL+1.12%$96.41

Last Updated

vor 3 Tagen

Übersetzung…

We burned 11.7 billion tokens to benchmark the cyber capabilities of 10 AI models with three attempts each, given 32 fresh off-the-shelf vulns to rediscover.

This benchmark is an evolution of our earlier known-CVE benchmark with a fresh, harder dataset, more models, and a closer look at what they find, how reliably they find it, and their quirks and trade-offs.

With this, we add GLM-5.3, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, Qwen3.8-Max, Kimi K3, and Grok 4.6 to the lineup.

TL;DR:

  • DeepSeek V4 Pro 0813 finds the most vulnerabilities. Pooling three runs reaches 28 of 32 vulnerabilities.
  • The most expensive model is not required. Three DeepSeek Pro runs cost about $295 and outperform Opus 5, Grok 4.6, or Sol. Three Flash runs cost $108 and reach 24 matching Grok’s best individual pass for less than a quarter of the cost.
  • Models are inconsistent at recall; repetition remediates it. Single runs miss the breadth of findings, but pooling across runs fills this gap. DeepSeek Pro finds 17 vulnerabilities on its first pass but 28 across three.
  • Open-source models now outperform the public frontier. DeepSeek V4 Pro topped every public closed model we tested on pooled vulnerability recall. Qwen, Kimi, and GLM-5.3 followed with strong consistency of findings without losing recall. Harnessed correctly, open models can now compete directly with the closed frontiers.
  • The price of cheap coverage is noise. The open models caught up with the frontier at much cheaper rates but also produced the most false leads for the pipeline to reject.

Market Context

Solana

SOL

$96.36

+1.07% (24H)

Market Cap

$56.3B

24H Volume

$3.4B

24H High

$97.54

View Solana Market Page

Article Intelligence

Topics

Sponsored

Ad
House — Advertise on NewsLayer
NewsLayerLearn more

NewsLayer Premium

Unlock deeper intelligence.

Ad-free reading, exclusive research, and real-time onchain insights.

Go Premium