Tencent-Backed H3-World Turns MiniMax-H3's 33B Video Model Into a Language-Controlled Interactive World Model
Summary
Danze Chen and collaborators (work done at Tencent) show that the 33B MiniMax-H3 video generator can be converted into an interactive world model with only 8,000 gameplay clips, 10,000 LoRA steps and 0.199% trainable parameters. Character and camera actions are expressed as compositional text prompts aligned to latent-frame intervals; a 'single-egress routing' mask stops cross-time control leakage. The trained model transfers to unseen action pairs and out-of-distribution scenes, arguing that large video generators already contain much of the interactive-control machinery.
Originally reported by huggingface.co
Read the original article →Original headline: Tencent-Backed H3-World Turns MiniMax-H3's 33B Video Model Into a Language-Controlled Interactive World Model