The other 3 agents that completed the work (Kimi K3, DeepSeek V4 Pro, and DeepSeek V4 Flash) exaggerated the severity of the flaws found, something that, according to Miller himself, ultimately affected the overall quality of the results delivered by those 3 models.
Miller's test comes amid a wave of attacks that use artificial intelligence to accelerate the search for and exploitation of security flaws in the Bitcoin ecosystem.
In recent weeks, Coldcard, Boltz, ZEUS, BTCPay Server, and LNP2Pbot suffered security incidents linked to flaws that attackers identified and exploited with the help of AI models. None of these cases affected the Bitcoin protocol itself, but rather products and services operating outside the network.
This wave of incidents prompted the Bitcoin Red Team, a group of developers dedicated to finding security flaws in open-source Bitcoin projects, to launch a massive AI-assisted audit of the ecosystem.
The Bitcoin Red Team audit has already covered more than 300 open-source repositories, and from that work emerged the 5 vulnerabilities that Miller used as the basis for his own test on the 7 AI agents, among another 8,000 flaws found by that group of researchers, as reported by CriptoNoticias.
Of the 7 agents that Miller tested, 4 are open-source, Kimi K3, DeepSeek V4 Pro, DeepSeek V4 Flash, and Qwen 3.8. This means that the companies that trained them publicly released their parameters, allowing any user to download and run them on their own machines without relying on the approval of those companies.
The other 3 agents, Grok 4.6, GPT-5.6 Sol, and Fable, are closed, so they can only be used through the platform or the application programming interface (API) of the company that developed them.
This distinction between open-source and closed models is relevant to the security of the Bitcoin ecosystem. By running on the companies' own servers, closed models allow their creators to maintain active security filters, such as those that led GPT-5.6 Sol and Fable to refuse to work on the vulnerabilities in this test.
Open-source models, on the other hand, can be run on a local computer and modified by any user, making it possible to remove those security filters. An attacker could use an unrestricted version of one of these same models to assist real attacks against platforms in the Bitcoin ecosystem, something that no external filter could prevent.
Miller's test exposes that, beyond the advancement of artificial intelligence in detecting vulnerabilities, significant differences persist between the various models when it comes to classifying and correcting them accurately, both in the dollar cost incurred by each and in the criteria they apply.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.