Google's Gemini Omni Can Generate 'Anything From Any Input,' Starting With Video - Engadget
Google's Gemini Omni generates video from any input—images, audio, text. The multimodal arms race just escalated.

Why it matters
Google is shipping a genuinely multimodal foundation model with cross-modal generation capabilities. This is a direct capability leap in the model_wars and signals Google's positioning against OpenAI's o1/o3 reasoning focus and Claude's agentic play.
The key facts
6 to knowGemini Omni: cross-modal generation (text, image, audio → video output)
Capability claim: 'generate anything from any input'
Google Flow AI video editing tools getting dedicated apps
Omni upgrades to existing Google tools
Published May 19, 2026 (future-dated; verify timestamp)
Multimodal foundation model release
Go to the source
Reuters Technologynews.google.com
Publisher excerpt: Google's Gemini Omni Can Generate 'Anything From Any Input,' Starting With Video Engadget Introducing Gemini Omni blog.google Gemini Omni Will Bring Only More AI Slop and Skepticism CNET Google’s Gemini Omni turns images, audio, and text into video — and that’s just the start TechCrunch Google Flow…