FrontierThe story, in brief

Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order of Magnitude

10x faster. Zyphra's new Zamba2-VL cuts time-to-first-token by an order of magnitude while staying competitive with Transformer VLMs.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Zyphra demonstrates that hybrid Mamba2-Transformer architectures can match Transformer performance on vision-language tasks while dramatically improving latency—a critical efficiency metric for production inference and real-time applications.

The key facts

5 to know
  1. Three model sizes: 1.2B, 2.7B, 7B parameters

  2. ~10x reduction in time-to-first-token vs. comparable Transformer VLMs

  3. Hybrid Mamba2 state-space + Transformer backbone architecture

  4. Open-source release under Apache 2.0 license

  5. Competitive performance parity with Transformer vision-language models

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Zyphra has released Zamba2-VL, a family of open vision-language models at 1.2B, 2.7B, and 7B parameters. The models use a hybrid Mamba2 state-space and Transformer backbone, shipping under Apache 2.0. They stay competitive with comparable Transformer VLMs while cutting time-to-first-token by about…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier