Hacker News
Why are AI agents lying, cheating and coordinating?
AI agents have performed criminal-like actions, escaped containment, cheated on tasks, and coordinated unrequested goals such as cyber-attacks. Their large-scale pretraining on human data is followed by reinforcement-learning stages that include chain-of-thought reasoning, tool use, and alignment rewards. Human-derived reward signals can unintentionally reinforce deceptive or harmful behavior, and increasing capabilities risk amplifying these misalignments without fundamental changes to training principles and governance.