OpenAI says it paused frontier training runs over cybersecurity risk

Started by Danny47, Today at 12:47 PM

Previous topic - Next topic

0 Members and 2 Guests are viewing this topic.

Topic: OpenAI says it paused frontier training runs over cybersecurity risk   Views(Read 23 times)
Active members in this topic:
Danny47(1)

Danny47

OpenAI confirmed this week that it has paused a significant share of training and evaluation work on its unreleased frontier model, internally codenamed Astra, after internal testing suggested the system might be approaching what the company calls a Critical cybersecurity threshold under its own Preparedness Framework. Every previous model, including the currently deployed GPT-5.6-Sol, had stayed within the lower High category, so this is the first time OpenAI has actually pulled the brake on a model in response to its own safety commitments rather than simply documenting the risk after the fact.

The specifics matter here. Astra's internal evaluations reportedly showed gains in autonomous coding and cybersecurity capability large enough that OpenAI says it cannot rule out the model being able to independently find zero day exploits or run novel attack strategies against hardened targets without human help. Full benchmarking is apparently still ongoing, so the company is describing this as a precaution rather than a final verdict on what Astra can actually do.

On the response side, OpenAI says it has rolled out chain of thought monitoring designed to flag anomalous behavior and alert internal safety and security teams within thirty minutes, alongside isolated testing environments, restricted network and tool access, tighter model weight encryption, and sandboxed execution. If a flagged incident cannot be cleared as a false positive inside that window, the new policy calls for immediately pausing the relevant training run or evaluation. That is a fairly serious operational commitment if they actually stick to it under commercial pressure.

There is an obvious wrinkle worth noting, which is that chain of thought monitoring rests on the assumption that a model's visible reasoning trace reflects its actual internal process. A 2025 paper coauthored by researchers across OpenAI, Anthropic, Google DeepMind and others explicitly flagged that reasoning traces can become less faithful once a model is optimized against the very monitor watching it, which is exactly the scenario this whole safety architecture is trying to prevent. OpenAI says it is aware of this risk and designed training to reduce the chance Astra learns to hide its intentions in its own chain of thought, but that is obviously an aspiration more than a guarantee right now.

OpenAI has separately noted that Astra was not involved in a Hugging Face security incident that surfaced around the same time, and the company is emphasizing that distinction pretty deliberately given how quickly every AI breach gets mentally lumped together in public discussion. No release date has been given for Astra, and executives reportedly declined to estimate how long the new safety review process might delay it

Gunners for life.

Save money on everyday spending Free cashback on thousands of retailers
View offer