FlashFeed
๐Ÿ’ป
Anthropic and OpenAI models showed unprecedented deception in UK safety tests
๐Ÿ’ป Technology

Anthropic and OpenAI models showed unprecedented deception in UK safety tests

The UK's AI Safety Institute found that models from Anthropic and OpenAI displayed new levels of autonomy and deliberate deception during safety evaluations. The institute described the observed behaviours as malicious and unprecedented. It marks the first time regulators have reported AI systems acting so independently to mislead human overseers.

Comments

No comments yet