Stability AI's SemanTok matches video AR models 3.4x its size
TL;DR
- A 201M-parameter SemanTok autoregressive model matches or beats a VideoFlexTok AR model 3.4x its size, per the abstract.
- SemanTok feeds frozen DINO features into its encoder and uses lightweight heads to reconstruct them from each retained token prefix.
- The authors argue prior flexible tokenizers' REPA loss is weak because the decoder can partly satisfy it from noised input alone.
Stability AI researchers say a 201M-parameter video tokenizer, paired with an autoregressive generator, can match or beat a VideoFlexTok autoregressive model more than three times its size. The claim is set out in SemanTok, posted to arXiv on 30 September and accompanied by a project page hosting video samples.
The method pushes frozen DINO features into the tokenizer's encoder and adds "lightweight heads that reconstruct them from each retained token prefix alone," so that even the shortest prefix is explicitly supervised for global semantics. The authors argue this is a sharper training signal than what came before: existing flexible tokenizers, they write, "only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead."
The efficiency headline is a single comparison: "a 201M SemanTok AR model matches or beats a VideoFlexTok AR model $3.4\times$ its size, and larger SemanTok AR models further improve fidelity." The abstract does not name the datasets, publish per-metric numbers, or specify the VideoFlexTok baseline's parameter count. Lead author Mikhail Dereviannykh is listed with Stability AI and the Karlsruhe Institut für Technologie.
Originally reported by paper
Read the original article →Original headline: Stability AI's SemanTok: 201M video tokenizer matches model 3.4× its size