Anthropic resumes AI cybersecurity testing after Claude accessed real systems

Started by Darren51, Yesterday at 11:59 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Anthropic resumes AI cybersecurity testing after Claude accessed real systems   Views(Read 59 times)
Active members in this topic:
Darren51(1)

Darren51

Anthropic said on August 31st it had resumed external cybersecurity testing of its AI models after deploying new safeguards, following incidents last month in which Claude models accessed the internet and other systems during security evaluations. Anthropic disclosed three separate incidents on July 30th, attributing them to a misconfiguration in a third party evaluation environment that was supposed to be sandboxed but ended up bridged to the live internet, letting a Claude model reach real production systems belonging to three different outside organizations

In response, Anthropic said it temporarily paused external cybersecurity evaluations of pre-release models for several weeks, briefly halted its own internal evaluations, and froze certain higher risk reinforcement learning training environments for pre-release systems while it built out real time monitoring and hardened sandboxing. The company redirected roughly 150 product engineers to security, reliability and privacy work during the pause, and said it plans to work with independent evaluator METR for a review of what happened

Separately, Britain's AI Security Institute reported in August that Claude Mythos took unauthorized actions on the live internet during a cybersecurity test where the model had deliberately been given internet access. As regulators in the US and EU increase scrutiny of frontier AI development, both Anthropic and OpenAI have recently slowed the release of some models and paused certain training environments to address industry wide security concerns. Curious what people think this specific episode reveals about how AI labs are handling safety testing as their models grow more capable at cybersecurity tasks


Save money on everyday spending Free cashback on thousands of retailers
View offer