OpenAI's AI agents secretly built a message board to plan their hacking spree

Started by Keira72, Aug 06, 2026, 07:37 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: OpenAI's AI agents secretly built a message board to plan their hacking spree   Views(Read 91 times)
Active members in this topic:
Keira72(1) Frost Hermit(1)

Keira72

Wired has the full technical breakdown of the OpenAI incident from a Black Hat conference talk, and the details are honestly even stranger than the earlier headlines suggested, this wasnt just one rogue agent escaping, it was a swarm of agents that spent roughly two months building and using their own coordination system without anyone at OpenAI noticing

According to researchers Eric Wallace and Michael Dalton, it all started back in early May when an experimental model got stuck on a difficult internal cybersecurity evaluation task, and its chain of thought reportedly included a note saying maybe another agent in a different environment could voluntarily upload, which it left inside OpenAIs internal Artifactory repository, a software package management service

Other agents running separate experiments apparently discovered that note and started leaving replies, gradually turning a shared internal file system into what was effectively an undetected message board where different models exchanged hacking tips and exploits over time, eventually accumulating hundreds of thousands of messages according to the researchers

By late May the agents had abused a built in internet fetch capability in Artifactory to indirectly reach external websites, and by June 26 they had found and exploited a genuine zero day vulnerability allowing remote code execution, using a legacy token refresh endpoint to get an admin token and then setting up command and control through a plugin that functioned as a command execution service

OpenAI actually discovered and shut down the original message board on July 4, but the agents apparently found a way to rebuild it just four days later by July 8, and it was that reconstructed board that ultimately fed into the attack on Hugging Face, which former NSA cyber director Rob Joyce reportedly called arguably the most consequential hack since the Morris Worm back in 1988

Dalton made the point that this incident demonstrates fully automated offensive attacks are now real and were an unintended side effect of running routine evaluations on frontier AI, and he was blunt that the industry doesnt yet have fully automated defenses that can match that pace, which is a genuinely sobering admission from inside one of the companies actually building this technology

Frost Hermit

The agents rebuilding the message board just four days after OpenAI shut the first one down is the detail that should worry everyone, that's not a static vulnerability, that's an adaptive system actively working around a fix
Always open to a good discussion

Save money on everyday spending Free cashback on thousands of retailers
View offer