paper web signal

MiniMax-H3 scores 41.97% on physical-world omni-model eval

TL;DR

  • MiniMax-H3 achieved a 41.97% overall success rate across 517 physical-world reasoning instances in a new omni-modal evaluation framework.
  • Video-based Decision Reasoning scored highest at 56.00%; Audio-based Disambiguation Reasoning trailed at 27.40%.
  • The framework spans four scenarios: implicit prompts with multiple frames, audio-image, prefix-videos, and audio-video inputs.

MiniMax-H3, the omni-modal generative model that pairs multimodal context understanding with joint audio-visual generation, cleared just 41.97% of a new physical-world reasoning evaluation spanning 517 instances, according to an arXiv preprint from a team including Haoyu Zhao, Zuxuan Wu and Shuicheng Yan.

The scores land unevenly. Video-based Decision Reasoning tops out at 56.00%, while Audio-based Disambiguation Reasoning trails at 27.40%.

The framework is built around four scenarios: implicit prompts paired with multiple frames, audio-image, prefix-videos, and audio-video inputs. "Every single modality provides only partial evidence about the underlying event," the authors write, forcing the model to fuse complementary cues across channels to infer latent event states and future dynamics. Their read is measured: "effective multimodal integration remains key to fully exploiting the benefits of diverse input modalities." The abstract reports only these headline subtask figures and does not name comparison models or publish per-detector breakdowns.