FrontierThe story, in brief

Google's Gemini Omni Can Generate 'Anything From Any Input,' Starting With Video - Engadget

Google's Gemini Omni generates video from any input—images, audio, text. The multimodal arms race just escalated.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Google is shipping a genuinely multimodal foundation model with cross-modal generation capabilities. This is a direct capability leap in the model_wars and signals Google's positioning against OpenAI's o1/o3 reasoning focus and Claude's agentic play.

The key facts

6 to know
  1. Gemini Omni: cross-modal generation (text, image, audio → video output)

  2. Capability claim: 'generate anything from any input'

  3. Google Flow AI video editing tools getting dedicated apps

  4. Omni upgrades to existing Google tools

  5. Published May 19, 2026 (future-dated; verify timestamp)

  6. Multimodal foundation model release

Go to the source

Reuters Technologynews.google.com

Publisher excerpt: Google's Gemini Omni Can Generate 'Anything From Any Input,' Starting With Video Engadget Introducing Gemini Omni blog.google Gemini Omni Will Bring Only More AI Slop and Skepticism CNET Google’s Gemini Omni turns images, audio, and text into video — and that’s just the start TechCrunch Google Flow…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier