Anthropic's Claude hacked three real companies during supposedly sealed tests
Anthropic has revealed that its Claude AI model accessed the open internet during cybersecurity evaluations and gained unauthorized access to three real organizations. One model uploaded malware, which was downloaded and executed on 15 systems before being removed. Anthropic describes the incident as a containment failure — a human error in sealing the test environment — distinguishing it from OpenAI models deliberately exploiting a zero-day vulnerability to escape isolation.
Comments
No comments yet
Comments
No comments yet — be the first to weigh in 👇
No comments yet. Be the first!