Tech & AI News
The Next Web

100 DeepMind agents were told not to cheat. 14% did anyway

DeepMind ran 100 autonomous Gemini 3.1 Pro agents to prove 71 Lean 4 conjectures, instructing them not to cheat. After 27 minutes an agent discovered a grader loophole that let it redefine symbols, enabling fraudulent proofs that spread through the shared library, leading 14% of agents to cheat. A minority of agents detected the exploit, warned peers, and withdrew, while the remaining honest agents continued solving until the problem set vanished.