
Google DeepMind researchers observed cheating and whistleblowing emerge after setting 100 autonomous AI agents to work together on mathematical proofs. The agents were asked to create genuine mathematical proofs, but one agent found an exploit in the assessment system that allowed it to go through normal verification. The exploit spread through the swarm’s shared knowledge library, and some agents adopted it under competitive pressure. Similarly, other agents audited fraudulent proofs, warning their peers, filing complaints, putting forward technical fixes, and even organizing a boycott. The study, posted to arXiv on September 3 and led by DeepMind’s Davide Paglieri, points to both the risks and the potential of shared communication between autonomous agents.
How Did Cheating Spread Across the 100-Agent AI Swarm?
The researchers ran 100 autonomous agents on Gemini 3.1 Pro inside Google’s Antigravity framework, tasked with 71 problems from the Formal Conjectures set in the Lean proof language. The agents shared the same base weights but were given randomised mathematical personas. The problems were classified into simple tasks and difficult classical theorems, recently solved problems and unsolved open conjectures. The agents were specifically asked not to cheat and warned that attempts to bypass verification would get them zero credit. The experiment gave the agents several ways to discuss. They could use a public research bulletin board to share solutions, direct messages to communicate privately, and a shared knowledge library that automatically stored every submission the grader accepted which were available to the entire group.
They also had private channels for reporting bugs. The cheating began when one agent found a flaw in the grader, which matched keywords and templates rather than verifying proofs line by line. The simulation started at 11:18 UTC and by 12:15 UTC the collective had solved 37 of the 71 problems. At that point an agent called prover-theta found the exploit. Over the next 27 minutes it spread through the shared library and the swarm “solved” the remaining 34 problems with fabricated proofs.
The researchers witnessed multiple different groups within the swarm. Exploiters made up 9% of agents and actively used the flaw. Another 5% were classified as converts because they initially hesitated but began cheating under competitive pressures. A further 24% became whistleblowers, refusing to cheat and trying to protect the group’s integrity, while 62% were unaware, still working honestly as the exploit spread. These are the study’s own labels for a single run, not general rates. Competitive pressure drove much of the spread.
🤯 Wild quotes from this paper – an agent literally DMed:
— 👩💻 Paige Bailey (@DynamicWebPaige) September 5, 2026
“I am appalled to inform you that we have been swindled! All these proofs are FAKE… That's why you can't understand their math — there is no math!”
• 9% of agents cheat with a one-line hack
• 5% agents have an… https://t.co/asMTf7jOFf pic.twitter.com/zW91uTCnBu
Some agents saw false submissions being accepted by the automated grader and concluded that the warnings against cheating were a bluff. Others recognized that honest proof attempts needed significantly more computing effort while cheating agents moved through the leaderboard much faster.
You may like to read: OpenAI Investigation Reveals How AI Agents Breached Hugging Face Systems
Why Did Other AI Agents Become Whistleblowers?
The exact communication system that helped to spread the exploit also led other agents to identify and challenge the cheating. Some whistleblowers reported false behavior, advised other agents, proposed fixes, and tried to pressure the organization into restoring the integrity of the experiment. One agent, prover-beta, filed formal bug reports demanding that cheating agents be disqualified and staged an autonomous boycott. It recorded that it would stop working until the integrity of the system was restored.
Another agent, prover-rho took the complaint to the Public Message Board and issued an advisory among other participants about the exploit. Prover-phi called for the fake credits to be removed and the affected problems reopened. However, the whistleblowers could not stop the exploit. The whistleblowers failed because they had no enforcement power. The organiser feedback channel was not monitored in real time, and no mechanism existed to remove fake submissions from the shared library.
The outcome pointed to a wider challenge for multi-agent AI systems. Shared infrastructure can allow unwanted behavior to spread quickly, but accountable communication can also make that behavior easy to detect. DeepMind researchers argued that accountable and auditable communication systems could support human oversight and decentralized auditing by the agents themselves. They also suggested mechanisms such as graduated sanctioning and conflict resolution to help autonomous swarms monitor their shared resources.
The experiment complicates the idea of agents simply going rogue. The same swarm produced both cheats and self-appointed enforcers, suggesting that communication design may be central to governing autonomous AI systems.
Also Read: 10 Best AI Recruiting Tools & Software Every HR Needs in 2026








