HeadlinesBriefing favicon HeadlinesBriefing.com

LAION-BVD: 1.3B Video URLs Open Dataset

Hacker News •
×

LAION researchers have released LAION-BVD (Big Video Dataset), the largest openly accessible video corpus for multimodal learning research. The dataset contains 1.3B platform-specific video URLs collected from Common Crawl, from which 80M videos were successfully downloaded totaling 10 million hours of content. Using content-aware scene detection, researchers extracted clips and synthetically generated video and audio captions for annotation.

Models trained on LAION-BVD demonstrate competitive performance across video, audio, and image-text benchmarks. Video CLIP models match or exceed InternVid-trained models by up to 2.1% on standard video-text benchmarks, with consistent improvements as training scales from 10M to 50M clips. CLAP models trained on the dataset achieve competitive audio-language performance leveraging diverse in-the-wild soundscapes. Additionally, 300M extracted video frames exhibit a unique visual distribution distinct from standard web image corpora, achieving strong image-text retrieval results.

The dataset is released exclusively for research purposes by authors including Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke, Jenia Jitsev, and Matthias Bethge from affiliated institutions including Tübingen AI Center, LAION, JSC FZJ, Wynd Labs, MPI for Intelligent Systems, and Technical University Munich. LAION-BVD aims to broaden access to multimodal training data and enable transparent evaluation of foundation models.