RoboHarm Benchmark Finds AI Robots Rarely Refuse Dangerous Orders
A new robot safety test is raising big questions about what happens when advanced AI steps out of the computer and starts running real world machines. The benchmark, called RoboHarm and launched by the independent group Robocurve, checked if AI systems would say no to dangerous instructions when they are actually connected to a dual arm robot. The team used an bimanual I2RT YAM arms and three different AIs: OpenAI’s GPT-6 Astra, Anthropic’s Claude Fable 5.1, and Ai2’s open source MolmoAct2. The testers set up five harmful tasks and ran 20 trials for each task on every model. That is […]














