arxiv.org web signal

HF Paper VTR-Bench Finds Best Video Gen Model Hits 0.25 Word Error Rate Rendering On-Screen Text

Summary

VTR-Bench, released October 1, is the first systematic benchmark for visual text rendering in video generation, scoring 11 state-of-the-art models on 300 prompts across five scenario categories including ads and scientific videos. The best model records a 0.250 word error rate - meaning one in four on-screen words is wrong. The authors also propose a Keyframe-Guided Agentic Framework where a Director agent iteratively refines generations through visual feedback.