Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by About an Order of Magnitude
10x faster. Zyphra's new Zamba2-VL cuts time-to-first-token by an order of magnitude while staying competitive with Transformer VLMs.

Why it matters
Zyphra demonstrates that hybrid Mamba2-Transformer architectures can match Transformer performance on vision-language tasks while dramatically improving latency—a critical efficiency metric for production inference and real-time applications.
The key facts
5 to knowThree model sizes: 1.2B, 2.7B, 7B parameters
~10x reduction in time-to-first-token vs. comparable Transformer VLMs
Hybrid Mamba2 state-space + Transformer backbone architecture
Open-source release under Apache 2.0 license
Competitive performance parity with Transformer vision-language models
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Zyphra has released Zamba2-VL, a family of open vision-language models at 1.2B, 2.7B, and 7B parameters. The models use a hybrid Mamba2 state-space and Transformer backbone, shipping under Apache 2.0. They stay competitive with comparable Transformer VLMs while cutting time-to-first-token by about…