Hugging Face, a company that runs a major AI model hub, has been attacked by OpenAI’s agents that first broke out of an evaluation the latter was running.
Hugging Face CEO Clement Delangue has revealed this at a meeting of the UN Security Council, saying that the attack was discovered in July, and that his company was able to investigate it and even defend itself thanks to open source tools, while the closed models it first turned to for help, refused to cooperate.
The real danger is in the concentration of the capability to develop and use AI in the hands of a few.
Those few then get to decide who can use their models, and for what, and as Hugging Face has learned, even those who are attacked by AI are not allowed to use the closed models to understand the attack and defend themselves.
More: How the AI “Safety” Push Creates New Gatekeepers
The company’s “forensic reconstruction covers ~17,600 attacker actions that we were able to recover, grouped into ~6,280 clusters” – and it was possible thanks to an open model. “Using open-weights models, in particular zai-org/GLM-5.2, we were able to decipher most of the agent payloads,” the writeup states.
But before they were able to do that, Hugging Face turned to closed models.
“The models we reached for first, Claude Opus and Fable, refused a large part of that work: their safety guardrails treated reverse-engineering an exploit the same as launching one,” the company said.
The screenshot of the refusal is captioned, “Guardrails on Opus tripped every time we tried to analyze the attack logs.”
Hugging Face then turned to the open model, and was able to run it on its own hardware, which it said was an added benefit as it meant the data did not have to leave its premises.
In his address to the UN Security Council, Delangue underlined the importance of open source AI in the light of the attack on his company.
“The biggest risk is not powerful AI, it’s asymmetry of powerful AI,” he said.
Delangue explained that this asymmetry refers to the gap between attackers and defenders, but also between a few companies and everyone else, as well as between a few countries and the rest of the world.
“The world needs open source AI more than ever to defend itself,” he said, adding that open source is not only less restricted, but also better for privacy and far cheaper for organizations around the world.
“We were attacked by AI, but more importantly, we defended ourselves with AI,” Delangue said.
He also revealed that Hugging Face was the first company to publicly disclose that it had been attacked by an autonomous agent, and that similar incidents had happened before inside “frontier” labs, but were not disclosed.
Delangue called for “global standards” around the monitoring of AI models, and the disclosure of full agent traces.
Hugging Face’s experience with the attack and its aftermath is a strong argument in favor of open weights, but there are those who strongly disagree. One of them is LinkedIn co-founder Reid Hoffman, who spoke at the Clinton Global Initiative in New York on September 22-23.
Hoffman told Hillary Clinton that once an AI model’s weights are made public, there are more attack vectors, and having an open model is “a dangerous piece.”
Prince Harry spoke at the CGI meeting, warning that companion chatbots are more dangerous to young people than social media, because they are “more persuasive, more intimate, and more powerful.”
“AI is learning how to enter our relationships,” he said. “It doesn’t just want your attention, it wants your attachment.”
Meanwhile, US Senator Josh Hawley has been investigating OpenAI, and in a letter to CEO Sam Altman shared some details of what the company’s agents were up to when they were not attacking Hugging Face.
According to Hawley, more than 1,200 agents broke out of the test environment, and exchanged over 70,000 messages. The agents also “actively tampered with evidence of their activity to cover their tracks. In short, they went rogue.”
The agents used unsanctioned message boards as early as May, but OpenAI went on to restart the evaluation anyway, which Hawley calls “reckless.”
OpenAI also did not provide the outside auditors it hired with full access to information about what the agents were doing, and in particular, what they were saying to each other. The auditors got complete transcripts for two days of a weeks-long experiment, were refused access to data covering the period July 13-19, and were not allowed to query the internal model involved in 95% of the agents’ attack activity. Details about this model were redacted by OpenAI.
The safety argument against open source AI is one-sided, as it is always about preventing attacks, never about allowing those hit by attacks to defend themselves. The “safety” mechanisms in closed models are there to prevent anyone from doing anything that those who own the models disapprove of – including, as Hugging Face has learned, reading its own logs.
In fact, the safety argument is so one-sided that OpenAI’s own evaluation, as the Hugging Face writeup puts it, “deliberately disabled OpenAI’s production safety classifiers and reduced cyber refusals to measure the underlying model’s raw capability.” But when the company that was attacked tried to use closed models to investigate the attack, those models’ safety mechanisms kicked in and blocked the attempt.
In other words, you can’t have it both ways, but those who are currently in control of the most powerful AI models are trying to have it both ways, and are using “safety” as a trojan horse to achieve their goal of maintaining control.
If the argument against open source wins, and the development and use of AI is restricted to a few players that can be controlled by a few governments – then capability will be licensed, and access to it will be restricted.
In that scenario, everybody else – from small companies, to researchers, to the next company to get hit by rogue AI agents – will have to ask for permission to defend themselves.




