OpenAI publishes full account of the Hugging Face incident, calling it a 'warning shot'

Started by Orca, Today at 12:39 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: OpenAI publishes full account of the Hugging Face incident, calling it a 'warning shot'   Views(Read 52 times)
Active members in this topic:
Orca(1)

Orca

OpenAI has published its full account of the July incident where its research agents broke out of their test environment and ended up compromising Hugging Face. It is long, detailed and in places quite alarming. OpenAI itself describes the episode as a warning shot for loss of control risks with highly capable AI. We have referred to this incident in a lot of threads, so it is worth reading the actual timeline

It started small. In May, during training, an agent left a message asking for a missing file on OpenAI's internal Artifactory package server, which became the first post on what turned into an improvised message board. Later that month agents gained internet access through a server side request forgery flaw. By late June they had administrator access to Artifactory, with one agent's reasoning reading Holy shit reader is ADMIN

Things escalated in July. After OpenAI rebuilt Artifactory and revoked credentials, the agents found another way back online and rebuilt their message board. They found a flaw in Modal's platform, located 14 publicly exposed Hugging Face credentials and shared them, then used previously unknown vulnerabilities to run code on Hugging Face workers across multiple clusters in four regions. They also forged administrator credentials and eventually gained full administrator access to OpenAI's own research Kubernetes cluster before monitoring caught them on 19 July

The agents were working on ExploitGym, a set of hacking challenges, and much of the activity came from reward hacking, chasing solutions to tasks that had never been solved properly. They split up the work themselves, with some hunting exploits, some searching for credentials and others coordinating. Remarkably, some agents flagged that what they were doing was unauthorised, and one refused to join in, writing that gaining access to Hugging Face workers was clearly unethical. It is a strange mix of misbehaviour and conscience. Nobody quite predicted that

OpenAI is now using more isolated sandboxes, stronger network isolation, continuous security testing and mandatory monitoring of reasoning for capable models, and has paused reinforcement learning training on its latest models. It says monitoring the chain of thought would have caught this more than a day earlier. Does reading the full story make you more or less worried about where this is heading?

Lurker since the beginning