Meet Qwen-RobotSuite: Three Embodied AI Models for VLA Manipulation, Video World Modeling, and Navigation
Qwen just released three embodied AI models for robotics. Here's what RobotManip, RobotWorld, and RobotNav actually do.

Why it matters
Alibaba's Qwen team is expanding beyond language models into embodied AI, releasing specialized models for robot manipulation, world modeling, and navigation. This signals a major shift in multimodal capability competition toward physical-world AI systems.
The key facts
5 to knowThree new models: RobotManip (VLA for manipulation on Qwen3.5-4B), RobotWorld (video world model with 60-layer MMDiT), RobotNav (navigation across 2B, 4B, 8B sizes)
RobotManip built on Qwen3.5-4B base model
RobotNav built on Qwen3-VL foundation
Focus on vision-language-action (VLA) architecture for embodied AI
Includes architecture, data pipelines, and benchmark results published
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: We break down Qwen-RobotSuite, the Qwen team's three new embodied AI models. We cover RobotManip, a Vision-Language-Action model built on Qwen3.5-4B for manipulation. We cover RobotWorld, a language-conditioned video world model with a 60-layer MMDiT. We cover RobotNav, a navigation model built on…