paper web signal

EgoTools benchmark: Gemini-3.1-Pro hits 51.7% on tool grounding

TL;DR

  • Gemini-3.1-Pro scores 66.9% overall on EgoTools-Bench but only 51.7% on Perception & Grounding, versus 83.3% for human experts.
  • Full supervised fine-tuning lifts Qwen3-VL-8B-Instruct from 50.0% to 60.9% on the 1,000-question benchmark under strict source-video separation.
  • EgoTools pairs roughly 100 hours of egocentric video with 1,000 eight-choice questions across four tracks: affordance, grounding, procedure, and spatial reasoning.

Gemini-3.1-Pro answers 66.9% of the questions on EgoTools-Bench correctly overall. On the Perception & Grounding track, the same model drops to 51.7%. Human experts score 83.3% on that track and 83.2% overall.

The benchmark is a 1,000-question, eight-choice evaluation built on roughly 100 hours of first-person recordings, split across four reasoning tracks: Affordance & Causality, Perception & Grounding, Procedural Dynamics, and Spatial Reasoning. The paper's framing is flat: "current multimodal video models remain limited in this form of tool-centric embodied reasoning."

Fine-tuning helps, but not enough to catch Gemini. "Full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9%, under strict source-video separation," the authors write. That still sits well short of the human overall. The project's GitHub repository labels the work [EMNLP26]; the retrieved arxiv page does not state acceptance.