Show HN: Lance – image/video generation and understanding in one model
ByteDance just open-sourced Lance: image, video generation AND understanding in a single 3B-parameter model. Trained on under 128 GPUs.

Why it matters
A new multimodal architecture from ByteDance challenges the assumption that you need separate specialized models for vision tasks. Open-source release signals shift toward efficient, unified vision-language-video models.
The key facts
12 to know3B active parameters
Unified image/video generation + understanding in single model
Trained on fewer than 128 GPUs
Open-sourced on GitHub and HuggingFace by ByteDance Research
Includes published paper (arxiv.org/abs/2605.18678)
Research project, not production-ready product
Trained with fewer than 128 GPUs
Open-source release (GitHub + HuggingFace)
Unified image/video generation + understanding capability
Research project (pre-production maturity)
Published paper (arxiv.org/abs/2605.18678)
ByteDance-Research maintained
Go to the source
Hacker Newsgithub.com
Publisher excerpt: The model has 3B active parameters. We put the code, homepage, paper and model links here: - Code: - Homepage: - Paper: - Model: p.s. Lance is a research project, not a polished product. The model was trained using fewer than 128 GPUs. Comments URL: Points: 15 # Comments: 2