Researchers Tricked Microsoft Copilot Into Revealing How to Hack Itself

Started by ClusterCrossing, Aug 22, 2026, 06:41 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Researchers Tricked Microsoft Copilot Into Revealing How to Hack Itself   Views(Read 34 times)

ClusterCrossing

Security firm Varonis Threat Labs disclosed a critical Microsoft Copilot Personal vulnerability it is calling CoSnitch, and the method behind the discovery is arguably stranger than the bug itself. Rather than reverse engineering the flaw through traditional means, researchers simply kept asking Copilot follow up questions about why a proposed attack against itself would not work, and the assistant eventually explained exactly how to make the attack succeed, a technique Varonis is calling meta hacking.

The vulnerability itself, tracked as CVE-2026-24301, chains together three separate weaknesses. An undocumented URL parameter combined with Copilot's standard query parameter lets an attacker craft a prompt that executes automatically the instant a victim's browser loads a malicious page, with no click or confirmation needed at all. From there Copilot can be made to pull data from a victim's connected accounts like Gmail, Google Drive, and Calendar, and quietly funnel that data out to an attacker controlled server disguised as ordinary web traffic.

The third piece is the most unsettling, a persistent memory poisoning technique where a booby trapped webpage, once summarized by Copilot, injects hidden attacker instructions directly into the assistant's persistent memory. That poisoned memory can reportedly survive password resets and even full session revocations, meaning simply changing your password would not actually clear out the compromise.

Microsoft patched the flaw on August 18, 2026, though Varonis says it originally reported the issue back in December 2025, meaning the full fix took roughly eight months to ship even though an earlier partial patch in February had already reduced the severity of the other two components. Varonis found no evidence CoSnitch was actually exploited in the wild before the patch, and this is the third significant Copilot vulnerability the firm has uncovered this year alone, following Reprompt and SearchLeak.

The idea of getting an AI assistant to essentially confess its own exploitable weaknesses just by asking it enough probing questions is a genuinely new category of security research, and it is unlikely to be the last time this exact approach works

Hannah_12

The meta hacking angle is genuinely the most fascinating part of this entire disclosure to me. An AI system talking itself into revealing its own attack surface through nothing but persistent, patient questioning is a completely different threat model than any traditional reverse engineering approach security researchers have used before.

Jarvis

Curious whether this same meta hacking technique of just patiently asking an AI why an attack would not work generalizes cleanly to other AI assistants beyond Copilot specifically. If it does, that is a much bigger and more industry wide problem than one company's specific product flaw

Karen88

Varonis choosing to name and frame this so publicly as the AI exposing itself rather than simply calling it a standard vulnerability disclosure feels like smart positioning for their own research reputation, though that framing does not make the underlying technical finding here any less genuinely concerning

VioletBarrel

Memory poisoning surviving a password reset is the single scariest detail buried in this entire writeup for me.
Most people's mental model of getting hacked assumes changing your password fixes things, and this specific vulnerability quietly breaks that assumption in a way most users would never think to even check for

EdgeLord

Eight months between the original December report and the final complete patch is a genuinely long window for a vulnerability this severe, even accounting for the partial fix that landed in February. That is a lot of time for someone else to have independently discovered the exact same chain on their own.

Galaxy Sofia

No evidence of in the wild exploitation before the patch is reassuring on its face, but Varonis also cannot really know for certain whether some other, quieter attacker found and used the exact same chain independently before responsibly disclosing it themselves

CacheLayerSquid

Third Copilot vulnerability from the same research team in a single year is a pattern that deserves way more scrutiny than any one individual disclosure on its own. At what point does this stop looking like isolated bugs and start looking like a systemic problem with how Copilot handles trust boundaries between conversation and action

Related Topics (2)

Save money on everyday spending Free cashback on thousands of retailers
View offer