Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation
Ant Group's 6B open-source robot model outperforms closed alternatives on cross-embodiment manipulation. Here's why VLAs just became commoditized.

Why it matters
Ant Group's LingBot-VLA 2.0 represents a significant capability jump in open-source vision-language-action models, trained on 60K hours of real robot data across 20 configurations. For robotics companies, this shifts the economics: a performant 6B model under Apache 2.0 means less dependency on closed APIs and faster iteration cycles.
The key facts
8 to know6B parameter model, Apache-2.0 licensed
Trained on 60,000 hours of data: 50,000 hours robot trajectories + 10,000 hours egocentric human video
Covers 20 robot configurations (arms, dexterous hands, waists, heads, mobile bases)
55-dimensional canonical action space for cross-embodiment generalization
Outperforms π0.5 and LingBot-VLA-1.0 on GM-100 generalist benchmark
Uses Mixture-of-Experts action expert (token-level, no load-balancing loss)
Dual-query distillation from LingBot-Depth and DINO-Video for geometric and temporal supervision
Released by Ant Group's Robbyant
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Ant Group's Robbyant has released LingBot-VLA 2.0, an Apache-2.0 vision-language-action model for cross-embodiment robot manipulation. The 6B checkpoint is pretrained on roughly 60,000 hours of data, spanning 50,000 hours of robot trajectories across 20 robot configurations and 10,000 hours of…