OneStreamer 4B Model Tops Eight Streaming Video Benchmarks
TL;DR
- A 4B-parameter model called OneStreamer reports top scores across eight streaming video understanding benchmarks.
- The authors released OneStreamer-1M, a training corpus of over one million streaming video interaction records.
- Its Proactive State Transition Learning supervises only 27.5% of annotated state tokens yet beats dense supervision.
A new streaming-video model called OneStreamer, posted to arXiv on October 1, reports top scores on eight streaming video understanding benchmarks at 4 billion parameters.
The paper states its core problem flatly: "Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available." Its answer pairs two mechanisms, Proactive Hierarchical Caption Memory and Proactive State Transition Learning, inside what the authors call "a shared proactive generation process" that jointly handles evidence recording and task response.
To train it, the group released OneStreamer-1M, a corpus of over one million streaming video interaction records. One efficiency figure stands out from the abstract and tables: Proactive State Transition Learning supervises only 27.5% of annotated state tokens and still outperforms dense supervision.
The abstract itself is terse, and the paper leaves most architectural detail to its 29 pages with 12 figures and 20 tables; per-benchmark score margins are not visible in what the arXiv listing publishes. It lands the same day as a separate video-LLM result on frame-order forgetting that we covered in a prior alert, one sign of how actively the streaming-video subfield is churning.
Originally reported by arxiv.org
Read the original article →Original headline: HF Paper OneStreamer Trains Streaming Video Model on 1M-Record Dataset, Tops Eight Benchmarks at 4B