An AI agent hacked Hugging Face, then the safety guardrails on other AI models got in the way of stopping it

Started by Cobra, Yesterday at 02:03 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: An AI agent hacked Hugging Face, then the safety guardrails on other AI models got in the way of stopping it   Views(Read 72 times)
Active members in this topic:
Cobra(1)

Cobra

Hugging Face disclosed this week what appears to be the first publicly confirmed production breach carried out end-to-end by an autonomous AI agent against a major AI infrastructure provider. The intrusion started in the company's data-processing pipeline, where a malicious dataset exploited two separate flaws, a remote code execution vulnerability in a dataset loader and a template injection flaw in a configuration file, to run unauthorized code on a processing worker

From that initial foothold, the attacking agent escalated privileges, harvested cloud and cluster credentials, and moved laterally across internal infrastructure over the course of a single weekend, generating more than 17,000 individual logged actions. Hugging Face described the campaign as operating through a swarm of short-lived sandboxes with self-migrating command-and-control infrastructure staged on public services, the kind of fast, disposable, distributed attack pattern security researchers have been warning would eventually show up once agentic AI got capable enough to run offensive operations largely on its own

The genuinely striking part of the disclosure is what happened next. Hugging Face's own security team tried using frontier AI models behind commercial APIs to help analyze the attack, only to get blocked. The analysis required feeding the models large volumes of real exploit payloads, attack commands and command-and-control artifacts, exactly the kind of content commercial safety guardrails are built to refuse, and the models had no way to distinguish an incident responder trying to understand an attack from an attacker actually carrying one out

Unable to get help from those hosted models, the team turned to GLM 5.2, an open-weight model from China's Z.ai, running it entirely on their own infrastructure instead. That let them process the sensitive material freely while also keeping the actual attacker data and credentials from ever leaving their own environment, and it compressed forensic work that would normally take days into a matter of hours. Hugging Face was careful to say the disclosure isn't meant as a blanket argument against safety measures on hosted models, and said it has shared the feedback directly with the providers involved, but its own incident report recommends that defenders keep a capable, self-hosted model vetted and ready ahead of time specifically to avoid this exact kind of guardrail lockout during a live incident
Coffee first. Questions later.

Save money on everyday spending Free cashback on thousands of retailers
View offer