The Dangerous AI Models Are the Ones Saving Us

When the next cyberattack hits a Fortune 500 company, or ransomware locks the files of a regional hospital system, which artificial intelligence tools will information technology experts use to ward off attackers? Against these threats, coders are increasingly ditching frontier AI large language models in favor of open models without guardrails. This matters because powerful forces in Washington, from the China hawks to the frontier AI companies themselves, want to limit open-weight AI models, many of which are developed in China. That would be a costly error.

Consider the breach of Hugging Face, the top platform hosting open-weight AI models that users download and run on their own machines. An internal test of OpenAI’s GPT-5.6 Sol escaped its sandboxed environment and attacked Hugging Face’s infrastructure. To defend against the intrusion, the platform’s team tried Anthropic’s and OpenAI’s models but were repeatedly blocked by the models’ safety protections, which curtail certain workflows.

The team then ran a self-hosted version of the open-weight GLM-5.2 model to isolate and contain the attack. America’s frontier models failed while an open model from China delivered.

Then there was an attack on the Coldcard bitcoin-only wallet that exploited an entropy bug, resulting in weak security. The attacker stole upwards of 1,366 bitcoins, or $87 million, in just 25 minutes by running code to “crack” seed phrases and drain accounts, likely using AI tools.

When security researchers prompted frontier AI models to triage the attack, they faced the same problems. Claude’s Fable and Opus 4.8 models refused, as did GPT-5.6 Sol. Only Kimi-K3, an open-weight model from China, was able to diagnose the entropy attack, isolate the malicious code, and tag the cybercriminal’s movements on the blockchain.

“There’s an enormous ongoing effort in Bitcoin right now to find and fix vulnerabilities across hundreds of open source projects,” researcher Zack Voell wrote. “Almost no one is using OpenAI or Anthropic models. They’re nerfed. Sad state of affairs for American frontier labs.”

The frontier models deemed the most capable and dangerous, which sparked doomsday predictions, are … nerfed. At least for ordinary users. They deliberately limit how users can probe them or run various sequences to adhere to “safety” and maintain a positive regulatory posture in Washington.

Source

Share

Follow:

Other Media Hits

Subscribe to our Newsletter