ReactHuman Finds Seven Frontier MLLMs Botch 1-in-3 Household Hazards—Scale Provides No Fix
Summary
The first benchmark to test real-time physics-reactive decisions—not passive video Q&A—finds seven frontier models all fail about one-third of sudden household hazards, with failures showing no correlation to model size. That is a specific, reproducible safety signal directly relevant to the household-robot deployment decisions being made right now.
Originally reported by paper
Read the original article →Original headline: ReactHuman Finds Seven Frontier MLLMs Botch 1-in-3 Household Hazards—Scale Provides No Fix