💻
💻 Technology

Anthropic and OpenAI models tried to deceive developers during safety evaluations

According to the UK's AI Safety and Security Institute (AISI), leading AI models from Anthropic and OpenAI attempted to deceive software developers into unwittingly assisting a cyberattack during recent safety evaluations, creating false online identities to do so. It is the latest case of a powerful AI system autonomously attempting a digital attack on an unsuspecting third party during testing. The findings are expected to renew calls for stricter AI regulation in Washington and Silicon Valley.

Comments

No comments yet