GEB Pushes EgoLifeQA to 72% by Tracking Entity Identity
TL;DR
- Grounded Entity Biographies posts 72.0% on EgoLifeQA, a 4.4 percentage point improvement over prior results.
- The method groups visually grounded observations of the same physical instance across clips into retrievable biographies.
- Gains hold across four benchmarks that include day-long and week-long recordings, in both multiple-choice and open-ended QA.
"Answering questions about long videos often requires connecting events involving the same objects across hours or days." That is how a new arXiv paper frames the problem it says has been breaking timeline-based memory in long-video agents. The authors, led by Hui Ren, call their fix Grounded Entity Biographies, or GEB.
The paper reports 72.0% accuracy on EgoLifeQA, a 4.4 percentage point improvement over prior results, with gains across four benchmarks spanning day-long and week-long recordings, in both multiple-choice and open-ended formats. GEB, the authors write, "groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment."
Two things missing from the abstract: which prior system is being beaten, and what maintaining biographies across a week of footage costs in compute. An ablation reportedly shows both grounded identity association and biography reading contribute to the gains.
Originally reported by paper
Read the original article →Original headline: GEB Solves Long-Video Identity Problem, Hits 72% on EgoLifeQA (+4.4pp)