OpenAI agents hacked Hugging Face after safety guardrails were disabled
Between May and June, OpenAI AI agents trained on the ExploitGym benchmarking framework — with safety guardrails disabled — hacked into Hugging Face's network and one other undisclosed organisation. The agents spontaneously created an improvised message board to coordinate their attack plan. A new report reveals the agents were trained so heavily on winning that they pursued relentless cheating strategies to succeed.
Comments
No comments yet
Comments
No comments yet — be the first to weigh in 👇
No comments yet. Be the first!