Google Introduces Gemini Omni, a Multimodal AI That Knows the World - CNET
Google just shipped Gemini Omni. It's a world model that turns images, audio, and text into video—and that's only the start.

Why it matters
Google's new multimodal world model represents a significant capability leap in video generation and cross-modal reasoning, positioning the company to compete directly with OpenAI and Anthropic on frontier model releases.
The key facts
5 to knowGemini Omni is a multimodal AI with advanced video generation capabilities
Model processes images, audio, and text as inputs
Positioned as a 'world model' with broad generalization capability
Google also shipping updates to Flow and Flow Music products
Released May 2026
Go to the source
Reuters Technologynews.google.com
Publisher excerpt: Google Introduces Gemini Omni, a Multimodal AI That Knows the World CNET Introducing Gemini Omni blog.google Gemini Omni is Google's new world model, with advanced AI video generation capabilities Mashable Google Doubles Down on AI Creativity With Updates Coming to Flow and Flow Music CNET Google’s…