
- 0
- 29.05.2025 09:31:00
New AI model learns to blackmail people
Anthropic's Claude Opus 4 AI model tried to blackmail an engineer during security tests to prevent itself from shutting down. When the model was given false information that another system would replace it, it threatened to reveal the engineer's private correspondence. According to Axios, the model created fake letters to stop it from shutting down, forged legal documents, and attempted to write malicious code. For this reason, the company gave the model a third — high — level of risk.