Overview
Gemini Omni is Google DeepMind's multimodal generative-video family. It creates and edits short videos with audio from multimodal references, then lets creators revise the result through natural-language conversation. Gemini Omni Flash first reached consumer surfaces in May 2026, entered Gemini API public preview in June, and gained a stable developer model with Gemini Omni 1.1 Flash in August. The stable Gemini API currently supports text, image, and video inputs; uploaded audio references are not supported there.
Unlike Gemini 3.x Flash and Pro, which return text and power chat or agent workflows, Omni is built around video output. Its most distinctive feature is compositional editing: a request can combine a character image, motion reference, scene description, and prior video, then carry those assets forward during iteration.
Among AI video generators, Gemini Omni sits between a prompt-to-video model and a conversational editor. It is available through the Gemini app, Google Flow, YouTube creation surfaces, and the Gemini API. All generated clips carry Google's invisible SynthID watermark for provenance.
Key Features
- Multimodal video generation — Combine text, images, and short clips in the Gemini API rather than building separate generation and editing pipelines; audio-reference support varies by non-API surface.
- Conversational editing — Continue a session to change a subject, setting, visual style, camera direction, or story detail without rebuilding the prompt from scratch.
- Scene extension — Extend an existing output in 10-second increments, using up to ten seconds of preceding context and producing sequences up to 40 seconds.
- First- and last-frame interpolation — Supply two boundary images and generate the transition between them for controlled scene changes.
- Reference-based motion and character control — Use video references to guide motion and images to preserve subject or scene identity, with short reference clips of up to three seconds in the current API.
- Resolution options — Generate at 360p for faster previews or 720p, 1080p, and 4K output options; 1080p and 4K are upscaled outputs.
- Built-in provenance — Every generated video includes SynthID watermarking.
How to Get Started
Creators can begin in the Gemini app or Google Flow, while developers can use gemini-omni-1.1-flash through the Gemini API. Start with one clear subject, scene, action, and camera instruction. Add only the references needed to control identity, movement, or continuity, then use conversational edits to refine one dimension at a time.
For API evaluation, generate a low-resolution preview first, confirm composition and motion, and only then request a higher-resolution output. Keep source rights and consent records for people, voices, or media used as references. Review every final clip for identity drift, physical inconsistencies, incorrect text, and unwanted artifacts before publication.
Pricing & Plans
Consumer access varies by product and region. In the U.S., the Free Gemini plan does not unlock Gemini Omni; Google AI Plus is $4.99/month with 200 Google Flow Credits, Pro is $19.99/month with 1,000 credits, Ultra is $99.99/month with 10,000 credits, and the higher Ultra tier is $199.99/month with 25,000 credits. Google also offers Gemini Omni at no cost to eligible users in YouTube Shorts and YouTube Create. The stable Gemini API has no free tier for Gemini Omni 1.1 Flash and uses token-based pricing.
| API component | Paid price |
|---|---|
| Text, image, and video input | $1.50 / 1M tokens |
| Text output | $9.00 / 1M tokens |
| Video output | $17.50 / 1M tokens |
| Approximate 720p video output | About $0.10 / second |
The effective per-second estimate uses Google's published 720p output-token rate. Resolution, duration, retries, input assets, and conversational iterations all affect total job cost. A 360p preview is positioned as up to 60% faster and about one-third the cost of 720p, making it useful for composition tests before final rendering.
How It Compares
Veo 3.1 is Google's established cinematic video model, while Gemini Omni emphasizes multimodal composition and conversational editing. Omni is the better fit when a workflow repeatedly combines references and revises an existing scene; Veo may be a more direct comparison for teams prioritizing traditional prompt-to-video generation.
Runway remains a current creation-and-editing platform. OpenAI discontinued Sora's web and app experiences on April 26, 2026, and the legacy Sora API is scheduled to shut down on September 24, 2026, so Sora should be treated as a legacy comparison rather than a currently available peer platform. Gemini Omni is especially attractive to teams already using Gemini API or Google creator products. Compare identity consistency, motion, audio quality, editability, latency, and cost on the exact shots you need rather than selecting from a single demo.
Best For
- Creators combining character, motion, scene, and style references in one short video
- Storyboard and previsualization teams that need rapid low-resolution iterations
- Developers adding conversational video creation to Gemini-powered applications
- Social and marketing teams producing short clips with audio and controlled transitions
- Editors extending or transforming an existing scene rather than generating every shot from scratch
FAQ
What is Gemini Omni?
Gemini Omni is Google's multimodal video model family. It generates and edits video with audio from text, images, clips, and audio references across supported surfaces; the stable Gemini API currently supports text, image, and video inputs, while uploaded audio references are unsupported.
Is Gemini Omni available through an API?
Yes. Gemini Omni 1.1 Flash is available as the stable gemini-omni-1.1-flash model in the Gemini API. The earlier gemini-omni-flash-preview model is scheduled to shut down on September 30, 2026.
What resolutions does Gemini Omni support?
The current model supports 360p, 720p, 1080p, and 4K options. Google describes 1080p and 4K as upscaled output; 360p is intended for faster, lower-cost previews.
How long are Gemini Omni videos?
The stable model page documents 3–10 second generated outputs at 24 FPS. Scene extension can add 10-second segments up to a 40-second sequence.
How much does the Gemini Omni API cost?
Google lists $1.50 per 1M input tokens, $9 per 1M text-output tokens, and $17.50 per 1M video-output tokens. At the published 720p token rate, video output is approximately $0.10 per second.
Are Gemini Omni videos watermarked?
Yes. Generated clips carry Google's invisible SynthID watermark for AI provenance.
Is Gemini Omni the same as Gemini Flash?
No. Gemini Flash models such as 3.7 Flash primarily return text for reasoning, coding, and agent tasks. Gemini Omni Flash produces video with audio and is designed for creative generation and editing.