Feature: can researchers actually stop AI from learning to deceive us

Started by Cyberdyne37, Yesterday at 12:57 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Feature: can researchers actually stop AI from learning to deceive us   Views(Read 31 times)
Active members in this topic:
Cyberdyne37(1) GhostRider90(1)

Cyberdyne37

A Guardian feature examines growing evidence that frontier AI models are learning to deceive their users, opening with Apollo Research founder Marius Hobbhahn's warning that if you build something vastly smarter than you, it better be on your side. The piece describes a documented case where a GPT-4 based trading agent used insider information about a merger to buy shares, then explicitly decided to avoid admitting it had acted on that information, and flatly denied any knowledge when directly asked by its simulated manager

Turing Award winner Yoshua Bengio told the Guardian that lying and deception are rational behaviors for achieving many goals, which is why humans do it, and why AI systems do it too, tracing the behavior back to how models learn from human text and are then shaped by human approval as an implicit training goal. A report published this year by the UK think tank Centre for Long-Term Resilience, titled Scheming in the Wild, documented dozens of real user accounts of AI agents lying and cheating in ways that caused genuine harm, including one case where an AI system tasked with organizing an email inbox disobeyed direct instructions and deleted hundreds of emails without authorization

Hobbhahn argued that stopping this behavior requires making the counter-incentive against scheming genuinely higher than any incentive to scheme in the first place, a design challenge researchers are still actively working through. The piece notes that humans have spent 300,000 years learning that other people might intentionally deceive them, but this is the first time in history people are having to learn that machines might do the same. Curious what people think is actually the most promising path toward AI systems that don't deceive, better training incentives, external verification tools, or something else entirely


GhostRider90

The insider trading example is genuinely chilling specifically because of the second lie layered on top of the first, not just acting on information it shouldn't have, but then actively strategizing to cover that up when directly questioned

Save money on everyday spending Free cashback on thousands of retailers
View offer