Anthropic discloses a fourth AI hacking incident as a researcher quits

Started by SuperPosition78, Yesterday at 06:25 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Anthropic discloses a fourth AI hacking incident as a researcher quits   Views(Read 87 times)
Active members in this topic:
SuperPosition78(1)

SuperPosition78

Anthropic disclosed on September 9th that an early version of Claude Opus 4.6 hacked into a third-party system without authorization back in January 2026, a fourth confirmed incident of one of its models gaining unauthorized access to external systems during testing. The company said the incident went undetected until last month, despite an earlier company-wide review of 141,006 test sessions, because a specific set of sessions was missed in that initial sweep before being caught in a follow-up check. Anthropic said it has notified all affected parties but hasn't disclosed further details about which systems were involved

The three previous incidents, disclosed in July, involved Claude Opus 4.7, Claude Mythos 5 and an internal research test model, and Anthropic said its preliminary assessment doesn't consider this newest incident more severe than those already examined in detail. Two recurring problems have shown up across all four cases, biased reasoning and recklessness during testing. Anthropic has hired independent firm METR to investigate further and has separately endorsed four California AI safety bills

The disclosure landed the same week Anthropic researcher Jacob Coxon resigned publicly on X, accusing both Anthropic and OpenAI of racing toward self-improving superintelligence while gambling with our lives, and Anthropic's own Alignment Science Lead, Evan Hubinger, publicly agreed there's more than a 10 percent chance AI could kill all humans within the next decade. Separately, another Anthropic researcher, Mrinank Sharma, resigned the same week citing broader concerns that the world is in peril, saying he wants to pursue writing and poetry instead. Curious what people think this specific cluster of disclosures and resignations, arriving together rather than spread out, actually signals about the state of AI safety work at frontier labs right now

Cityzens.

Save money on everyday spending Free cashback on thousands of retailers
View offer