paper web signal

Every frontier MLLM fails Hong Kong navigation test

TL;DR

  • Pedestrian collision rates ran 76.3% to 90% across every frontier MLLM walked through a Hong Kong sandbox.
  • Long-range point-to-point navigation dropped to 0-3.8% success even as short-range success hit 15-75%.
  • Nine model families were tested, including GPT-5.5, Claude Opus 5, Gemini 3.6-Flash, Doubao-Seed-2.0-Pro and Kimi-K3.

Pedestrian collision rates ran between 76.3% and 90% across every frontier vision-language model walked through a Hong Kong sandbox, according to a paper on arXiv introducing UrbanGround, "the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data."

The authors put nine MLLM families through a first-person walk-around: GPT-5.5, 5.4 and 5.2; Claude Opus 5 and Opus 4.6; Gemini 3.6-Flash and 3.1-Pro; Doubao-Seed-2.0-Pro; GLM-5V-Turbo; and Kimi-K3. Short scenes went reasonably well. Visual recognition landed between 77.7% and 93.8%, and short-range navigation reached 15% to 75%.

The problem starts once the destination gets farther away. Long-range point-to-point navigation success collapsed to a 0-3.8% band across the board. Orientation understanding sat at 23.3-58.3%, with several systems, the paper says, approaching "the random-guessing baseline."

The authors are blunt about the pattern: "local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction." When routes were blocked mid-task, "agents often continue to produce locally compliant movement without recovering safe goal-directed progress."

The abstract frames the shared symptom in one line: "orientation and pedestrian-aware movement remain unreliable." No single model in the nine-family lineup is called out as the exception.