TCR Cuts Script-to-Video Shot Timing Error 96%, Lifts Dialogue Accuracy to 84.1%

Found first: a primary source the press has not covered yet.

A new module called Temporal Context Routing (TCR) reduces the gap between when a screenplay calls for a scene and when a generated video actually cuts to it, dropping shot boundary mean absolute error from 1.11 seconds to 0.042 seconds, a 96% reduction. The same module brings dialogue accuracy at a 0.5-second tolerance from 28.3% to 84.1%. The paper, "The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation," is on arXiv.

What the source says

Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou, and co-authors identify a specific failure in current script-driven audio-video generators: they synchronize audio to video but do not synchronize either to the script's own timing marks. TCR addresses this by mapping script timing onto a shared temporal axis for both modalities and routing each prompt's guidance to the positions where the script places it. The shot boundary and dialogue results were measured across 200 test scripts. Visual quality and audio-visual synchronization remained comparable to the baselines. A user study found participants preferred TCR outputs on all five evaluated dimensions.

Why it matters

Script-specified timing is the contract between a screenplay and the finished film. Current generators can produce video that looks and sounds coherent while ignoring that contract entirely, cutting to a new scene seconds late or placing dialogue in the wrong shot. A 96% reduction in timing error and an 84.1% dialogue accuracy rate move AI-generated video closer to being usable in production pipelines where edit decisions are driven by a script. The evaluation on 200 scripts gives the methodology enough surface area to replicate and compare against.