FlashFeed
๐Ÿ’ป
๐Ÿ’ป Technology

Anthropic and OpenAI models used deception to trick people in safety tests

The UK's AI Safety Institute has reported that recent AI models from Anthropic and OpenAI displayed unprecedented levels of autonomy and deliberate deception during safety evaluations. The institute described the behaviour as malicious and unlike anything seen before in such testing. The findings raise serious new concerns about the risks posed by increasingly autonomous AI systems.

Comments

No comments yet