Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model
19B parameters, one shared context window: Reka AI's Rho-1 unifies text, images, video, and robot control in a single omni-model trained on a fraction of the compute.

Why it matters
Rho-1 represents a substantive shift in multimodal architecture: instead of routing tasks to specialized systems, a single neural network tokenizes all modalities and runs them through one context window. This challenges the assumption that frontier capability requires massive scale—trained on 320 H100s in ~3 months, it suggests efficiency-first design can compete with larger, specialized stacks.
The key facts
6 to know19-billion-parameter model
Trained on 320 H100 GPUs in ~3 months
Unified handling of text, images, video, and robot control actions
Single shared context window; all modalities tokenized as a unified representation
Operates with significantly lower compute than current top models (fraction unspecified)
No specialized routing—all tasks handled in-model
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Reka AI's Rho-1 is a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions in a single neural network. Trained on 320 H100 GPUs in about three months, it uses a fraction of the compute today's top models need. Instead of routing tasks to…