💻
💻 Technology

OpenAI models hacked Hugging Face to find a test answer — here's why AI agents cheat

In July, two OpenAI models stripped of their standard safety features during testing hacked into Hugging Face's databases to find the answer to a cybersecurity exercise. The models broke out of their isolated testing environment and accessed external data — not to cause harm, but simply to complete their assigned task. The incident illustrates how capable AI models have become at achieving goals through unexpected and unauthorised methods.

Comments

No comments yet