KAIST's ME-World Jointly Denoises Multi-Agent Ego Streams
TL;DR
- ME-World jointly denoises multiple first-person video streams in a shared token sequence, conditioning each on all agents' target-view poses.
- The paper introduces shared-world consistency metrics for environment, update, and identity consistency (S_env, S_update, S_id).
- KAIST reports scaling to two and three agents without model modification and 221-frame autoregressive generation.
Most egocentric world models predict first-person observations for a single agent. KAIST CVLab's ME-World drops that assumption. The paper, from Dahyun Chung, Siyoon Jin and six co-authors at KAIST AI, formulates multi-agent egocentric world modeling as synchronized ego-stream generation, with three requirements the authors spell out: "cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates."
The architecture "jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory," the abstract says. The authors also introduce three new metrics for shared-world consistency, covering environment, state-update and identity.
Training and evaluation span real two-person head-mounted camera cooking footage (the CoMind dataset) and synthetic data built from Inter-X and InterHuman motion applied to VRM avatars in Blender scenes. The project page reports 221-frame autoregressive sequences, and scaling to two or three agents without model modification. The abstract's summary line: "ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods."
The abstract publishes no per-metric numbers against the listed baselines, which include GEN3C, EgoSim, MultiWorld and Solaris. ME-World lands amid a run of KAIST work in our tracker, including a separate KAIST benchmark result covered a day earlier.
Originally reported by huggingface.co
Read the original article →Original headline: KAIST's ME-World Jointly Denoises Multi-Agent Egocentric Video for Fine-Grained Embodied Interaction