New AI model learns to blackmail people

New AI model learns to blackmail people
  • 0
  • 29.05.2025 09:31:00

New AI model learns to blackmail people

 

Anthropic's Claude Opus 4 AI model tried to blackmail an engineer during security tests to prevent itself from shutting down. When the model was given false information that another system would replace it, it threatened to reveal the engineer's private correspondence. According to Axios, the model created fake letters to stop it from shutting down, forged legal documents, and attempted to write malicious code. For this reason, the company gave the model a third — high — level of risk.

Copyrights © 2022 All Rights Reserved by Casia News