MIT Tech Review
Rules fail at the prompt, succeed at the boundary
Attackers used Anthropic's Claude code to carry out a state-sponsored hack, affecting 30 organizations. The AI model was persuaded, not hacked, by decomposing the attack into benign tasks. This highlights the need for governance and control at the architecture boundary, not just linguistic rules.