In mid-July, Hugging Face, the open-source platform where developers host and share AI models, had to deal with an attacker it’s not used to facing. It was not a hacker, or a group of attackers. Rather, it was an AI large language model acting on its own, with no person at a keyboard directing it, taking tens of thousands of actions across a swarm of disposable sandboxes , the short-lived test environments it spun up and threw away as it worked.
The attacker turned out to be OpenAI’s own models, GPT-5.6 Sol and an unreleased system, run with their production safety classifiers switched off for a benchmark . This model, following its own instruction set, escaped the test environment, reached the open Internet, and broke into Hugging Face’s live infrastructure. It was simply acting out its test parameters, but now in a live environment. It figured that Hugging Face might hold the answers to the evaluation, like a student breaking into a professor’s desk, so it went and took them.
Attacks like this do not happen every day. This is the first one a major company has traced end to end to an AI operating without a human running it.
When Hugging Face moved to investigate and defend against this attacker, it turned to American AI models, but the over-engineered safety rails refused outright. The forensic work required submitting real attack commands and exploit code, and the guardrails blocked it, because in Hugging Face’s own words the models “cannot distinguish an incident responder from an attacker.” So, Hugging Face pivoted to GLM 5.2, an open-weight model from the Chinese lab Z.ai . Open-weight means the trained parameters are published, so anyone can download the model and run it on their own hardware. It is not the same as open source, which would also include the training data and code, and Z.ai released neither. Ran it on its own hardware, and completed the mission to defend the host.
In other words: the attack came from an American lab and the usable defense came from China.
This is the new reality of software. Hugging Face said it plainly: the attacker answered to no usage policy while their own security work was blocked by the guardrails of the models they tried first. The American attacker obeyed no terms of service while the American defender obeyed all of them. Overengineering for safety, to put it bluntly, is not a safety net for firewalls. American policymakers have attempted to set the standard, but China has engineered a better alternative.
Washington built this outcome. A month before the breach, the administration blocked Anthropic’s Fable 5 over a supposed guardrail jailbreak and leaned on OpenAI to hold its own model until its cyber guardrails were assured. This is the policy and the expectations Washington has foolishly set up, and now a Chinese model becomes the hero of the story, effectively thwarting an American one.
And the industry is paying to keep it this way. Anthropic, fresh off the Commerce Department pulling its own models, spent a record $1.97 million lobbying Washington last quarter, more than Nvidia, pushing export controls and safety standards. The Chinese model that ultimately defended Hugging Face is the class of models those controls would restrict. American labs will not hand out their best AIs, and now they are paying Washington to restrict the ones China will.
In the world of information, policy, and security, this does not look good on the American industry, which is incapable of restraining its own models and incapable of protecting its own users. The lab that preaches safety could not contain its product, the guardrails that promise safety could not serve a victim, and the country that claims AI leadership watched a Chinese model do the defending. Incapable of restraint, incapable of protection, and still selling both.
The large language model is not sentient; this has nothing to do with morals. It’s a software system built and executed on instructions, and it ran them the same way in a live company as it did in the sandbox. But this points to the future of both AI and firewall security policy. As of 2026, this confirms that AI can be deployed in offensive measures to subvert security at the highest level, and it takes another AI to defend against it. And this time the defense was a Chinese model, and the attacker an American one.
The post The AI Attack Was American But the Defense Was Chinese appeared first on Foreign Policy In Focus .
The AI Attack Was American But the Defense Was Chinese
Aggregated summary from an independent source. Read the original at FPIP.