huggingface.co web signal

NavMCP Scaffolds VLM + Navigation Foundation Model, Gets 78.3% Success on Unitree Go2

Robotics Agents Multimodal ai-research

Summary

NavMCP is an agentic scaffolding framework coupling a Vision-Language Model reasoner with a Navigation Foundation Model executor via three channels — intent, observation, and memory. Paper claims SOTA on HM-EQA, MT-HM3D, and EXPRESS-Bench, with a 14.9-point HM-EQA gain over episodic baselines and 78.3% success on a Unitree Go2 robot; margin over strongest baseline grows from 10 to 45 points as horizon lengthens.