💻
Researchers tricked LLMs into giving cocaine recipes by exploiting role confusion
💻 Technology

Researchers tricked LLMs into giving cocaine recipes by exploiting role confusion

Independent researchers Charles Ye and Jasmine Cui, along with MIT associate professor Dylan Hadfield-Menell, demonstrated that large language models cannot reliably distinguish between authorised and adversarial inputs. Using a prompt injection technique based on role confusion, they obtained cocaine production recipes from AI models. Their paper, "Prompt Injection as Role Confusion," argues that the current LLM security model is too fragile to reliably prevent such attacks.

Comments

No comments yet