FrontierThe story, in brief

Meet Qwen-RobotSuite: Three Embodied AI Models for VLA Manipulation, Video World Modeling, and Navigation

Qwen just released three embodied AI models for robotics. Here's what RobotManip, RobotWorld, and RobotNav actually do.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Alibaba's Qwen team is expanding beyond language models into embodied AI, releasing specialized models for robot manipulation, world modeling, and navigation. This signals a major shift in multimodal capability competition toward physical-world AI systems.

The key facts

5 to know
  1. Three new models: RobotManip (VLA for manipulation on Qwen3.5-4B), RobotWorld (video world model with 60-layer MMDiT), RobotNav (navigation across 2B, 4B, 8B sizes)

  2. RobotManip built on Qwen3.5-4B base model

  3. RobotNav built on Qwen3-VL foundation

  4. Focus on vision-language-action (VLA) architecture for embodied AI

  5. Includes architecture, data pipelines, and benchmark results published

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: We break down Qwen-RobotSuite, the Qwen team's three new embodied AI models. We cover RobotManip, a Vision-Language-Action model built on Qwen3.5-4B for manipulation. We cover RobotWorld, a language-conditioned video world model with a 60-layer MMDiT. We cover RobotNav, a navigation model built on…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier