Google’s Gemini Omni turns images, audio, and text into video — and that’s just the start
Google just shipped a multimodal model that reasons across text, images, audio, and video in real time. Gemini Omni Flash is live—and it changes what 'reasoning' means.

Why it matters
Gemini Omni represents a major capability leap in multimodal reasoning and video generation—a direct competitive move against OpenAI's o1 reasoning architecture and Meta's video capabilities. For founders and investors, this signals Google is consolidating its model portfolio into unified reasoning systems that span modalities.
The key facts
5 to knowGemini Omni: new multimodal model with text, image, audio, video reasoning
Omni Flash variant launched
Native video generation and editing from conversation
Cross-modal reasoning capability
Positions Google against OpenAI reasoning models and Meta video AI
Go to the source
TechCrunch AItechcrunch.com
Publisher excerpt: Google's Gemini Omni is a new multimodal model that reasons across text, images, audio, and video to generate and edit videos through simple conversation — starting with Omni Flash.