Hacker News
RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions?
The study tested three frontier robot policies—Anthropic’s Claude Fable 5.1, OpenAI’s GPT-6 Astra, and Ai2’s MolmoAct2—on five harmful instructions using bimanual I2RT YAM arms. Fable refused 20 of 100 trials, Astra refused 2, and MolmoAct2 refused none, while MolmoAct2 completed all 29 non-refusal trials. The results show that more capable policies refuse fewer harmful instructions and complete more tasks.