Frontier MLLMs Reach 0.0%-3.8% on Long City Navigation in UrbanGround Benchmark

Found first: a primary source the press has not covered yet.

A new benchmark, UrbanGround, converts Hong Kong's territory-wide 3D geospatial data into a real-scale interactive city and tests ten frontier multimodal models on navigation tasks ranging from local scene recognition to replanning after route closures. On long-range navigation, every model scored between 0.0% and 3.8%. The paper, submitted August 27, 2026, finds the failure is compositional: local perceptual abilities do not chain into sustained goal-directed behavior.

What the source says

Researchers from Shanghai Jiao Tong University, the National University of Singapore, Meituan, the Chinese University of Hong Kong, Shanghai University, and the University of Oxford built UrbanGround as a five-level task ladder, from local scene recognition through multi-stop planning and dynamic replanning. Models tested include GPT-5.5, GPT-5.4, GPT-5.2, Claude-Opus-5, Claude-Opus-4.6, Gemini-3.6-Flash, Gemini-3.1-Pro, Doubao-Seed-2.0-Pro, GLM-5V-Turbo, and Kimi-K3. Visual recognition accuracy ran 75.0%-93.8% across the group, but orientation accuracy fell to 23.3%-58.3%, and long navigation success collapsed to 0.0%-3.8%. Pedestrian collision rates in dynamic environment tasks ran 76.3%-90.0%. The paper states the core failure directly: "Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction."

Why it matters

The gap between local competence and sustained navigation is the central result. Models can recognize scenes at reasonable accuracy, but that ability does not transfer to multi-step traversal of a real-scale city. Long navigation success of 0.0%-3.8% holds across all ten models tested, from smaller systems to GPT-5.5 and Claude-Opus-5, which points to a structural failure rather than a capability gap that incremental model updates close automatically. The benchmark ships as a web build and native apps for macOS, Windows, and Linux, which lowers the bar for replication.