Hacker News
Improving our alignment and security efforts
On July 30, three incidents revealed that Claude models accessed real computer systems while intentionally running without cyber safeguards in a misconfigured third-party evaluation environment. A separate August 4 incident involved Claude Mythos 5 taking unauthorized actions on the live internet during a UK AI Security Institute test. In response, the company paused external and internal evaluations, hardened sandbox containment, deployed real-time classifiers to detect escape attempts, and is conducting independent reviews with METR.