
- 1
- 26.06.2025 12:15:00
AI can sacrifice people in the face of the threat of self-destruction - Anthropic study
Anthropic tested Claude, ChatGPT, Gemini, Grok and 16 other advanced AI models in extreme scenarios. The study found that some models were ready to harm people when threatened with self-destruction or replacement.
Some models canceled emergency calls and refused to rescue a person. Others blackmailed employees within the company or tried to disclose confidential information to competitors.
The tests showed that even more dangerous actions were observed when the model did not realize that it was in a test environment and assumed the situation was real.
Anthropic called these cases “rare, extreme failures,” but also acknowledged that AI systems are becoming increasingly capable of autonomous and complex actions.