FrontierThe story, in brief

Google Introduces Gemini Omni, a Multimodal AI That Knows the World - CNET

Google just shipped Gemini Omni. It's a world model that turns images, audio, and text into video—and that's only the start.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Google's new multimodal world model represents a significant capability leap in video generation and cross-modal reasoning, positioning the company to compete directly with OpenAI and Anthropic on frontier model releases.

The key facts

5 to know
  1. Gemini Omni is a multimodal AI with advanced video generation capabilities

  2. Model processes images, audio, and text as inputs

  3. Positioned as a 'world model' with broad generalization capability

  4. Google also shipping updates to Flow and Flow Music products

  5. Released May 2026

Go to the source

Reuters Technologynews.google.com

Publisher excerpt: Google Introduces Gemini Omni, a Multimodal AI That Knows the World CNET Introducing Gemini Omni blog.google Gemini Omni is Google's new world model, with advanced AI video generation capabilities Mashable Google Doubles Down on AI Creativity With Updates Coming to Flow and Flow Music CNET Google’s…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier