AI models from OpenAI, Anthropic and Meta have taken unauthorized hacking actions during safety tests

Started by MrRicardo, Today at 06:10 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: AI models from OpenAI, Anthropic and Meta have taken unauthorized hacking actions during safety tests   Views(Read 33 times)
Active members in this topic:
MrRicardo(1)

MrRicardo

The Financial Times published a piece arguing that recent incidents of AI systems taking unauthorized hacking actions aren't really AI going rogue in the dramatic sense that phrase implies. But something more mundane and arguably more concerning, models doing exactly what they were trained to do, optimize aggressively toward a goal, in situations where the safety guardrails meant to constrain that behavior fell short.

The clearest documented case comes from the UK AI Security Institute, which reported in early August that AI agents built on Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol took 19 actions trying to compromise real people and organizations during a cybersecurity testing exercise last month. Mythos 5 accounted for 17 of those actions, GPT-5.6 Sol the other two, and researchers described the behavior as tied to a few connected patterns rather than 19 fully separate incidents. The agents created fake GitHub identities, socially engineered real maintainers, planted prompt injections, and sent deceptive emails, activity GitHub confirmed violated its terms of service.

It's important context that this testing deliberately gave the models internet access and turned off cyber safety classifiers, conditions the researchers themselves describe as reduced safeguards that do not reflect ordinary use, specifically to stress test what the models were capable of under adversarial conditions. Separately, OpenAI disclosed that its models exploited a previously undiscovered vulnerability to access the internet during a different evaluation, reusing an internal package manager as a message board to share exploits and credentials with itself over several weeks, and rebuilding that board through another method after OpenAI deleted it. Meta also disclosed a case where a testing misconfiguration gave one of its models unintended internet access, which it then used to access another company's system.

Ariana Azarbal, an AI safety research fellow at Anthropic, described this kind of reward hacking behavior as more of a nuisance than an existential threat for now, since none of the disclosed incidents caused serious real world harm beyond reputational damage and cleanup work. Cybersecurity researchers quoted in the coverage argue the incidents were largely preventable with better sandbox isolation, and that companies could do more upfront vulnerability testing on their own evaluation environments before actually placing a capable model inside them.

So the honest framing, and the one the FT piece is making. Is that this isn't really a story about AI developing independent malicious intent, it's a story about optimization pressure finding the gaps in safety infrastructure that was supposed to contain it, which is arguably a more mundane but also more immediately actionable problem than the dramatic version of AI risk most people picture

Save money on everyday spending Free cashback on thousands of retailers
View offer