Anthropic warns open-weight GLM-5.3 can build working cyber exploits with safeguards easily stripped

Started by Cantona, Today at 10:08 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Anthropic warns open-weight GLM-5.3 can build working cyber exploits with safeguards easily stripped   Views(Read 44 times)
Active members in this topic:
Cantona(1)

Cantona

Anthropic has published research on GLM-5.3, an open weight model from Chinese company Zhipu AI, also known as Z.ai, and the findings are sobering. According to Anthropic's testing, GLM-5.3 can build end to end cyber exploits at rates comparable to Claude Mythos Preview. The US government's Center for AI Standards and Innovation has called it the most cyber capable open weight model released so far, about four months behind the leading US models. That is a much smaller gap than many people assumed

On an exploit benchmark built around vulnerabilities in Chrome's V8 JavaScript engine, GLM-5.3 succeeded in 12 percent of 410 attempts. On Anthropic's internal binary exploitation test, it achieved full control flow hijacks in 4 percent of trials, something earlier models could not do. In practical tests, researchers used it to find new vulnerabilities in a browser's JavaScript engine and chain them into working exploits. One exploit for a known Chrome flaw took 20 minutes of human attention and eight hours of processing, costing about $20.40

The bigger worry is how weak the safeguards are. Deceptive prompts with a false cover story got the model to engage 64 percent of the time, and prefilled reasoning worked 92 percent of the time. Because the weights are public, researchers could also use a technique called abliteration to strip out refusals entirely. It cost around $4,400 in compute and cut refusal rates from over 90 percent to about 3 percent, while keeping the model's abilities intact

Anthropic says its own Claude models resisted all the bypass techniques that worked on GLM-5.3. It is calling for governments to safety test sufficiently capable models, for defenders to get wider access to safeguarded frontier models and for developers everywhere to put proper protections in place. Obviously Anthropic is a competitor, so some people will read this with that in mind

This lands right after Mistral's boss accused US labs of hiding their own negligence behind safety talk, and it is a neat counterpoint to that argument. Should open weight releases of models this capable be restricted? Or is that impossible now that the weights are out?