OpenAI finds 'self-replicating prompt injections' that spread like computer worms

Started by James95, Today at 10:05 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: OpenAI finds 'self-replicating prompt injections' that spread like computer worms   Views(Read 73 times)
Active members in this topic:
James95(1)

James95

The New Stack reports on a new OpenAI research paper describing what it calls self-replicating prompt injections. OpenAI says it found instances of its GPT models being susceptible to an AI version of a worm attack. A normal prompt injection tricks a model into doing something malicious. This kind goes a step further by also instructing the model to copy the injection into its outputs, so it spreads to other systems and agents

OpenAI identified three main variants. Email based injections arrive in a message and tell an agent to include the attack in outgoing emails. Filesystem based ones spread through files, with one example using a fake system warning to make a model delete reports while embedding the attack into other files. Multi-hop attacks chain several steps together, for example telling an agent to fetch more instructions from Slack, act on them and pass the payload along

The company found these using GPT-Red, a self-play training framework introduced in July that pits attacker models against defender models. To test for worm-like behaviour, researchers added a goal for the attacker to make the target repeat the injection on a public output channel. The tests used internal research versions based on GPT-5.4 mini and GPT-5.5

Importantly, OpenAI says it saw no real world incidents or impact, and everything happened in simulated tool calls during internal training and evaluation. The problem was discovered in June and disclosed publicly this month. OpenAI is now building self-reproduction goals into GPT-Red training so future models learn to resist these attacks

This is exactly the sort of risk that grows as agents get access to email, files and messaging tools. One compromised message could in theory spread through an organisation's agents without any human noticing. Should companies be connecting agents to email and Slack at all yet? And is it good that OpenAI is publishing this, or does it give attackers ideas?


Save money on everyday spending Free cashback on thousands of retailers
View offer