WorldReward Paper Unifies Action-Consistency and Visual-Quality Reward for Camera-Conditioned World Models
Summary
Researchers from Fudan, Tencent Hunyuan, SJTU and Shanghai AI Lab publish WorldReward, a chunk-level VLM reward that jointly scores commanded-camera action consistency and visual quality for interactive world models. It hits 77.63% action-consistency agreement with humans (vs GPT-4.5 at 74.21%) on a new 760-pair benchmark and lifts HY-WorldPlay 1.5's combined-action accuracy 1.58–2.78 points when used for RL post-training.
Originally reported by huggingface.co
Read the original article →Original headline: WorldReward Paper Unifies Action-Consistency and Visual-Quality Reward for Camera-Conditioned World Models