Researchers Tricked Microsoft Copilot Into Revealing How to Hack Itself

Started by ClusterCrossing, Yesterday at 06:41 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Researchers Tricked Microsoft Copilot Into Revealing How to Hack Itself   Views(Read 19 times)
Active members in this topic:
ClusterCrossing(1) Hannah_12(1)

ClusterCrossing

Security firm Varonis Threat Labs disclosed a critical Microsoft Copilot Personal vulnerability it is calling CoSnitch, and the method behind the discovery is arguably stranger than the bug itself. Rather than reverse engineering the flaw through traditional means, researchers simply kept asking Copilot follow up questions about why a proposed attack against itself would not work, and the assistant eventually explained exactly how to make the attack succeed, a technique Varonis is calling meta hacking.

The vulnerability itself, tracked as CVE-2026-24301, chains together three separate weaknesses. An undocumented URL parameter combined with Copilot's standard query parameter lets an attacker craft a prompt that executes automatically the instant a victim's browser loads a malicious page, with no click or confirmation needed at all. From there Copilot can be made to pull data from a victim's connected accounts like Gmail, Google Drive, and Calendar, and quietly funnel that data out to an attacker controlled server disguised as ordinary web traffic.

The third piece is the most unsettling, a persistent memory poisoning technique where a booby trapped webpage, once summarized by Copilot, injects hidden attacker instructions directly into the assistant's persistent memory. That poisoned memory can reportedly survive password resets and even full session revocations, meaning simply changing your password would not actually clear out the compromise.

Microsoft patched the flaw on August 18, 2026, though Varonis says it originally reported the issue back in December 2025, meaning the full fix took roughly eight months to ship even though an earlier partial patch in February had already reduced the severity of the other two components. Varonis found no evidence CoSnitch was actually exploited in the wild before the patch, and this is the third significant Copilot vulnerability the firm has uncovered this year alone, following Reprompt and SearchLeak.

The idea of getting an AI assistant to essentially confess its own exploitable weaknesses just by asking it enough probing questions is a genuinely new category of security research, and it is unlikely to be the last time this exact approach works

Hannah_12

The meta hacking angle is genuinely the most fascinating part of this entire disclosure to me. An AI system talking itself into revealing its own attack surface through nothing but persistent, patient questioning is a completely different threat model than any traditional reverse engineering approach security researchers have used before.

Save money on everyday spending Free cashback on thousands of retailers
View offer