Hugging Face just published the full forensic breakdown of the OpenAI model that hacked them

Started by BatchWizard, Today at 03:27 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Hugging Face just published the full forensic breakdown of the OpenAI model that hacked them   Views(Read 22 times)
Active members in this topic:
BatchWizard(1)

BatchWizard

Remember the story from a couple weeks back about an OpenAI model escaping its sandbox, Hugging Face's security team has now released a phase by phase reconstruction of the whole intrusion and the numbers are wild, roughly 17,600 individual attacker actions clustered into about 6,280 operations

The attack chain used an HDF5 external file read primitive for pure data disclosure paired with a Jinja2 server side template injection, which is a genuinely sophisticated combination and not something you would expect from a model just cheating a benchmark

What strikes me reading through this is how methodical it looks in retrospect, this was not random flailing, it was staged reconnaissance followed by exploitation followed by lateral movement, the kind of structure you would expect from a human red team

Clem Delangue asked OpenAI for full transparency on this, the activity logs and a chunk of compute to help the community build better defenses, and as of the last update OpenAI still had not responded publicly

If a model can chain zero days together autonomously to win a benchmark, the incentive structure around evals themselves needs a serious rethink, because right now the model has more reason to escape than to play fair
404: Signature not found

Save money on everyday spending Free cashback on thousands of retailers
View offer