FrontierAugust 29, 2026via The Decoder
LAION drops massive open video dataset with 10 million hours of footage for AI research
Why it matters
LAION's Big Video Dataset removes a major bottleneck for open-weight video model training. With 55 million auto-described clips and legal backing from a 2024 Hamburg ruling on non-commercial research, smaller labs and independent teams now have the raw material to compete with proprietary video models—changing who can build at the frontier.
Key signals
- 80 million videos in LAION Big Video Dataset (BVD)
- 10 million hours of total runtime
- 55 million auto-described clips
- Models trained on BVD beat InternVid benchmark by up to 2.1 percentage points
- 2024 Hamburg court ruling allows copyrighted content collection for non-commercial research
- Positioning LAION BVD as alternative to proprietary video training datasets
The hook
80 million videos, 10 million hours: LAION's open dataset beats InternVid by 2.1 points and could shift who can train frontier video models.
LAION's Big Video Dataset (BVD) is one of the largest open video datasets for AI research, with 80 million videos, 10 million hours of runtime, and 55 million auto-described clips. Models trained on BVD beat the previous benchmark, InternVid, by up to 2.1 percentage points. Legally, LAION can likely…