Google · 2026-05-19 · major
Gemini Omni — Google's 'Create Anything From Any Input' Model Ships With Conversational Video Generation
Gemini Omni Flash generates and edits video from text, image, audio, or video input through conversation. It reasons about physics and continuity, ships today in the Gemini app and on YouTube Shorts, and watermarks output with SynthID.

A single Gemini model that turns text, images, or audio into edited video on request.
What is it?
Gemini Omni is a new family of generative models from Google. The first release, Gemini Omni Flash, takes any input — text, images, audio, or video — and produces video, with image and audio output planned. It also edits existing clips when you describe the change in plain language.
How does it work?
Rather than bolting a separate text-to-video model onto Gemini, Omni folds generation into the model itself, so it draws on Gemini's reasoning to decide what should happen next in a scene. Google says it pairs an understanding of physics with world knowledge of history, science, and culture. Every output carries an imperceptible SynthID watermark that the Gemini app, Chrome, and Search can detect.
Why does it matter?
It puts conversational video creation and editing inside a mainstream app and on YouTube Shorts at no cost, lowering the bar for short-form video. The mandatory SynthID watermark is Google's answer to provenance concerns as synthetic video gets easier to make.
Who is it for?
creators, short-form video makers
Try it
Gemini app or YouTube Shorts (Gemini Omni Flash)