Ant Group’s Robbyant Unveils LingBot-VA 2.0: A Causal Video-Action Model Built Natively for Physical AI
Ant Group's LingBot-VA 2.0 hits 225 Hz control. Physical AI just got faster—and built natively for robots, not borrowed from video models.

Why it matters
Ant Group is advancing embodied AI with a purpose-built video-action foundation model designed specifically for physical robotics, signaling a shift away from adapting consumer AI models toward domain-specific architectures. This represents a meaningful capability leap in real-time robotic control and could influence how other labs approach Physical AI infrastructure.
The key facts
8 to knowLingBot-VA 2.0 is a causal video-action foundation model built natively for embodiment (not fine-tuned from video generators)
Achieves 225 Hz asynchronous control rate
Features Foresight Reasoning for predictive state modeling ahead of execution
Uses causal DiT (Diffusion Transformer) architecture
Implements sparse-MoE video stream processing
Includes semantic visual-action tokenizer
Published as technical report by Ant Group's Robbyant division
Re-grounds predictions on every real observation (closed-loop)
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Ant Group's Robbyant has released the LingBot-VA 2.0 technical report — a Physical AI video-action foundation model built from scratch for embodiment rather than fine-tuned from a video generator. It predicts future states ahead of execution through Foresight Reasoning, re-grounds on every real…