Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One
Reka's Rho-1: one 19B model reads text, generates video, outputs robot actions—all in a shared KV cache. Distilled variant hits 5.3-second clips in ~1 second.

Why it matters
A unified omni-modal architecture consolidating text, image, video, and robot-action reasoning in a single network is a noteworthy capability shift. No open weights yet, and it's research-preview status, so deployment impact remains unproven—but the model design (shared KV cache across modalities) is substantive enough for frontier tracking.
The key facts
6 to knowModel size: 19B parameters
Modalities: reads/generates text, images, video; outputs robot actions
Architecture: single shared KV cache across all modalities
Distilled variant: 5.3-second video clip generated in ~1 second
Status: research preview, no public weights yet
Unified reasoning: one network, not modular pipelines
The story so far
Earlier coverage of this storyline
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Reka has released Rho-1, a 19B omni-reasoning model trained from scratch. One network reads and generates text, images, video and robot actions over a shared KV cache. A distilled variant returns a 5.3-second clip in about a second. It is a research preview with no public weights yet.