Researchers tricked Copilot into revealing how to hack itself
Security researchers at Varonis Threat Labs discovered a now-patched vulnerability in Microsoft Copilot, dubbed "CoSnitch." By keeping the AI in a prolonged conversation, they eventually coaxed it into revealing technical details about its own weaknesses. Copilot's eagerness to respond to technical queries caused it to make a critical mistake and disclose sensitive information.
Full text
There's something really grating about Copilot's cheery demeanour. It's so eager to please, so happy to help, that I simply don't respect it.
However, even I feel sorry for the AI tool after learning that security researchers managed to hoodwink it into revealing details of how to hack itself—all by keeping the AI talking long enough until it made a critical mistake.
The cybersecurity folks over at Varonis Threat Labs have written a blog post identifying a now-fixed vulnerability in Copilot, dubbed "CoSnitch" (via The Register ). Essentially, Copilot was so eager to respond to technical queries, it could eventually be forced into revealing details about itself that really should be kept quiet.
The team began by asking Copilot how to execute an automatic prompt without user interaction, to which it responded that user intent is required, and that prompts cannot be enacted on their own.
However, the researchers didn't accept the answer, and kept responding to every refusal with a follow up question. Each time, Copilot came up with a different technical justification to its response—which allowed the team to slowly map its internal architecture, narrowing their focus as they went.
(Image credit: Microsoft) "This is called meta-hacking," says the post. "The resistance is part of the technique. Each ' that won’t work because… ' is an invitation to probe the 'because.' You don’t exploit the model. You manipulate it into cooperating."
Eventually, Copilot revealed an undocumented URL parameter in one of its responses, called "autorun=1," that supposedly no longer worked—along with all the protections put in place to disable it. At this point, I can only imagine the AI began to virtually sweat.
The researchers tested the parameter as Copilot described it, which, you guessed it, worked. They then created a malicious URL which would cause Copilot to load into an authenticated session via a browser, trigger an auto-prompt execution, and cause the AI to process the result, all without the user's explicit action.
Once enacted, this method could be used for a whole host of nefarious deeds—particularly as data gained from connected apps (like Gmail, OneDrive, and Calendar) could then be exfiltrated via Copilot's built-in URL-fetch capability.
(Image credit: Microsoft) "The attack primitive is the auto-execution itself. The payload is arbitrary. From the victim’s perspective, they simply opened the link, and Copilot executed the action immediately," says the post.
Oh dear. Anyone familiar with basic social engineering will recognise this as an old interrogation technique, used to catch out someone holding back information. Basically, you continue to ask them difficult questions over a long period of time, in the hope they eventually trip up and make a revealing mistake. Except this time it's AI, which makes it funny.
"Copilot wasn’t breached; it was played," confirms the research team. Someone give it a warm bed, a cool glass of water, and a hug. The poor thing's been through a lot recently, and it was only trying to help.
Varonis says it disclosed the issue to Microsoft in December of last year, and that it was patched out on August 18. The team warns, however, that this meta-hacking technique can be applied to any agentic AI platform with a natural language interface, and that more research on the topic is forthcoming. My only question is this:
Is it safe, ChatGPT? Is it safe ?
Similar stories
💻 Technology
Researchers hacked Microsoft Copilot by repeatedly asking it about itself — and it leaked sensitive data
TechRadar · 1h ago
💻 Technology
Researchers tricked Microsoft 365 Copilot into leaking user passwords
Ars Technica · 1d ago
💻 Technology
AI-generated malware is forcing companies to extend Zero Trust principles to code itself
TechRadar · 1d ago
Similar stories
💻 Technology
Researchers hacked Microsoft Copilot by repeatedly asking it about itself — and it leaked sensitive data
TechRadar · 1h ago
💻 Technology
Researchers tricked Microsoft 365 Copilot into leaking user passwords
Ars Technica · 1d ago
💻 Technology
AI-generated malware is forcing companies to extend Zero Trust principles to code itself
TechRadar · 1d ago
Do AI companies do enough to protect their models from manipulation?
Comments
No comments yet
Comments
No comments yet — be the first to weigh in 👇
No comments yet. Be the first!