FrontierThe story, in brief

Google’s Gemini Omni turns images, audio, and text into video — and that’s just the start

Google just shipped a multimodal model that reasons across text, images, audio, and video in real time. Gemini Omni Flash is live—and it changes what 'reasoning' means.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Gemini Omni represents a major capability leap in multimodal reasoning and video generation—a direct competitive move against OpenAI's o1 reasoning architecture and Meta's video capabilities. For founders and investors, this signals Google is consolidating its model portfolio into unified reasoning systems that span modalities.

The key facts

5 to know
  1. Gemini Omni: new multimodal model with text, image, audio, video reasoning

  2. Omni Flash variant launched

  3. Native video generation and editing from conversation

  4. Cross-modal reasoning capability

  5. Positions Google against OpenAI reasoning models and Meta video AI

Go to the source

TechCrunch AItechcrunch.com

Publisher excerpt: Google's Gemini Omni is a new multimodal model that reasons across text, images, audio, and video to generate and edit videos through simple conversation — starting with Omni Flash.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier