Found first: a primary source the press has not covered yet.
LAION has published LAION-BVD, an open video dataset built from 80 million downloaded videos totaling 10 million hours, the largest publicly available video corpus for multimodal pretraining. The dataset spans video, audio, and image modalities and is released under CC-BY-4.0.
What the source says
The team, led by first author Andreas Hochlehnert, sourced 1.3 billion platform-specific video URLs from CommonCrawl and successfully downloaded 80 million of them. Content-aware scene detection was used to extract clips, for which synthetic captions were generated covering both video and audio. Scene-changing frames were extracted separately as an alternative image-text data source, producing a visual distribution the authors describe as distinct from standard web image corpora. ViCLIP and CLAP models trained on LAION-BVD achieve competitive performance on standard video-text and audio-text benchmarks, with the paper reporting consistent improvements as training or model scale increases and comparing results against InternVid-trained baselines.
Why it matters
No prior open dataset approaches this scale of video content. Large-scale video-text corpora used by major AI labs have not been publicly released, keeping competitive multimodal video pretraining out of reach for most researchers. LAION-BVD addresses that gap directly. The CC-BY-4.0 license permits commercial use with attribution, a broader grant than a research-only release. That combination of scale and license terms sets this apart from earlier open video efforts.