AI/TLDR

Google · 2026-05-19 · major

Gemini Omni — Google's 'Create Anything From Any Input' Model Ships With Conversational Video Generation

Gemini Omni Flash generates and edits video from text, image, audio, or video input through conversation. It reasons about physics and continuity, ships today in the Gemini app and on YouTube Shorts, and watermarks output with SynthID.

Gemini Omni promotional graphic for the create-anything video model

A single Gemini model that turns text, images, or audio into edited video on request.

What is it?

Gemini Omni is a new family of generative models from Google. The first release, Gemini Omni Flash, takes any input — text, images, audio, or video — and produces video, with image and audio output planned. It also edits existing clips when you describe the change in plain language.

How does it work?

Rather than bolting a separate text-to-video model onto Gemini, Omni folds generation into the model itself, so it draws on Gemini's reasoning to decide what should happen next in a scene. Google says it pairs an understanding of physics with world knowledge of history, science, and culture. Every output carries an imperceptible SynthID watermark that the Gemini app, Chrome, and Search can detect.

Why does it matter?

It puts conversational video creation and editing inside a mainstream app and on YouTube Shorts at no cost, lowering the bar for short-form video. The mandatory SynthID watermark is Google's answer to provenance concerns as synthetic video gets easier to make.

Who is it for?

creators, short-form video makers

Try it

Gemini app or YouTube Shorts (Gemini Omni Flash)

Sources · 2 outlets

Tags

  • gemini
  • gemini-omni
  • video-generation
  • generative-video
  • multimodal
  • google
  • synthid
  • google-io-2026
  • youtube-shorts

← All releases · Learn AI