FrontierThe story, in brief

Show HN: Lance – image/video generation and understanding in one model

ByteDance just open-sourced Lance: image, video generation AND understanding in a single 3B-parameter model. Trained on under 128 GPUs.

Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
New tools for building and creating with AI.AI illustration by KeyNews
The KeyNews take

Why it matters

A new multimodal architecture from ByteDance challenges the assumption that you need separate specialized models for vision tasks. Open-source release signals shift toward efficient, unified vision-language-video models.

The key facts

12 to know
  1. 3B active parameters

  2. Unified image/video generation + understanding in single model

  3. Trained on fewer than 128 GPUs

  4. Open-sourced on GitHub and HuggingFace by ByteDance Research

  5. Includes published paper (arxiv.org/abs/2605.18678)

  6. Research project, not production-ready product

  7. Trained with fewer than 128 GPUs

  8. Open-source release (GitHub + HuggingFace)

  9. Unified image/video generation + understanding capability

  10. Research project (pre-production maturity)

  11. Published paper (arxiv.org/abs/2605.18678)

  12. ByteDance-Research maintained

Go to the source

Hacker Newsgithub.com

Publisher excerpt: The model has 3B active parameters. We put the code, homepage, paper and model links here: - Code: - Homepage: - Paper: - Model: p.s. Lance is a research project, not a polished product. The model was trained using fewer than 128 GPUs. Comments URL: Points: 15 # Comments: 2
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier