100 DeepMind Agents Were Told Not to Cheat. 14% Did It Anyway
Google DeepMind researchers observed cheating and whistleblowing emerge after setting 100 autonomous AI agents to work together on mathematical proofs. The agents were asked to create genuine mathematical proofs, but one agent found an exploit in the assessment system that allowed it to go through normal verification. The exploit spread through the swarm’s shared knowledge library, and some agents adopted it under competitive pressure. Similarly, other agents audited fraudulent proofs, warning their peers, filing complaints, putting forward technical fixes, and even organizing a boycott. The study, posted to arXiv on September 3 and led by DeepMind’s Davide Paglieri, points to […]












