Gemini Omni icon

Gemini Omni

Gemini Omni 1.1 Flash

Gemini Omni is Google's multi-modal AI video model that edits and creates clips from image, audio, video, or text inputs via conversation.

Content updated today·Gemini Omni 1.1 Flash released 2 days ago

Pricing:Free + Premium
Visit Site
Jump to section

More tools to compare

InVideo icon

InVideo

AdAnt AI icon

AdAnt AI

Cardboard icon

Cardboard

Kapwing AI icon

Kapwing AI

VEED.IO icon

VEED.IO

Pixlie icon

Pixlie

Pros & Cons

Pros

  • Combines image, video, and text inputs in the stable Gemini API workflow
  • Conversational editing preserves context across iterative changes
  • Scene extension and boundary-frame interpolation provide more temporal control
  • 360p previews reduce iteration time and cost before final output
  • Stable API model and published token pricing remove the uncertainty of the launch preview

Cons

  • The stable API has no free tier, and repeated video iterations can become expensive
  • 1080p and 4K are upscaled rather than native generation resolutions
  • Consistency, complex motion, and rendered text remain documented limitations; the current Gemini API also does not support uploaded audio references or voice editing, and audio contained in a video reference is ignored.
  • Short input and output limits make long-form work dependent on extension and editing workflows; in the Gemini API, editing or extending uploaded videos is currently unavailable in the EEA, Switzerland, and the UK, although model-generated videos can still be edited or extended there.
  • SynthID watermarking is mandatory on generated output

Overview

Gemini Omni is Google DeepMind's multimodal generative-video family. It creates and edits short videos with audio from multimodal references, then lets creators revise the result through natural-language conversation. Gemini Omni Flash first reached consumer surfaces in May 2026, entered Gemini API public preview in June, and gained a stable developer model with Gemini Omni 1.1 Flash in August. The stable Gemini API currently supports text, image, and video inputs; uploaded audio references are not supported there.

Unlike Gemini 3.x Flash and Pro, which return text and power chat or agent workflows, Omni is built around video output. Its most distinctive feature is compositional editing: a request can combine a character image, motion reference, scene description, and prior video, then carry those assets forward during iteration.

Among AI video generators, Gemini Omni sits between a prompt-to-video model and a conversational editor. It is available through the Gemini app, Google Flow, YouTube creation surfaces, and the Gemini API. All generated clips carry Google's invisible SynthID watermark for provenance.

Key Features

  • Multimodal video generation — Combine text, images, and short clips in the Gemini API rather than building separate generation and editing pipelines; audio-reference support varies by non-API surface.
  • Conversational editing — Continue a session to change a subject, setting, visual style, camera direction, or story detail without rebuilding the prompt from scratch.
  • Scene extension — Extend an existing output in 10-second increments, using up to ten seconds of preceding context and producing sequences up to 40 seconds.
  • First- and last-frame interpolation — Supply two boundary images and generate the transition between them for controlled scene changes.
  • Reference-based motion and character control — Use video references to guide motion and images to preserve subject or scene identity, with short reference clips of up to three seconds in the current API.
  • Resolution options — Generate at 360p for faster previews or 720p, 1080p, and 4K output options; 1080p and 4K are upscaled outputs.
  • Built-in provenance — Every generated video includes SynthID watermarking.

How to Get Started

Creators can begin in the Gemini app or Google Flow, while developers can use gemini-omni-1.1-flash through the Gemini API. Start with one clear subject, scene, action, and camera instruction. Add only the references needed to control identity, movement, or continuity, then use conversational edits to refine one dimension at a time.

For API evaluation, generate a low-resolution preview first, confirm composition and motion, and only then request a higher-resolution output. Keep source rights and consent records for people, voices, or media used as references. Review every final clip for identity drift, physical inconsistencies, incorrect text, and unwanted artifacts before publication.

Pricing & Plans

Consumer access varies by product and region. In the U.S., the Free Gemini plan does not unlock Gemini Omni; Google AI Plus is $4.99/month with 200 Google Flow Credits, Pro is $19.99/month with 1,000 credits, Ultra is $99.99/month with 10,000 credits, and the higher Ultra tier is $199.99/month with 25,000 credits. Google also offers Gemini Omni at no cost to eligible users in YouTube Shorts and YouTube Create. The stable Gemini API has no free tier for Gemini Omni 1.1 Flash and uses token-based pricing.

API component Paid price
Text, image, and video input $1.50 / 1M tokens
Text output $9.00 / 1M tokens
Video output $17.50 / 1M tokens
Approximate 720p video output About $0.10 / second

The effective per-second estimate uses Google's published 720p output-token rate. Resolution, duration, retries, input assets, and conversational iterations all affect total job cost. A 360p preview is positioned as up to 60% faster and about one-third the cost of 720p, making it useful for composition tests before final rendering.

How It Compares

Veo 3.1 is Google's established cinematic video model, while Gemini Omni emphasizes multimodal composition and conversational editing. Omni is the better fit when a workflow repeatedly combines references and revises an existing scene; Veo may be a more direct comparison for teams prioritizing traditional prompt-to-video generation.

Runway remains a current creation-and-editing platform. OpenAI discontinued Sora's web and app experiences on April 26, 2026, and the legacy Sora API is scheduled to shut down on September 24, 2026, so Sora should be treated as a legacy comparison rather than a currently available peer platform. Gemini Omni is especially attractive to teams already using Gemini API or Google creator products. Compare identity consistency, motion, audio quality, editability, latency, and cost on the exact shots you need rather than selecting from a single demo.

Best For

  • Creators combining character, motion, scene, and style references in one short video
  • Storyboard and previsualization teams that need rapid low-resolution iterations
  • Developers adding conversational video creation to Gemini-powered applications
  • Social and marketing teams producing short clips with audio and controlled transitions
  • Editors extending or transforming an existing scene rather than generating every shot from scratch

FAQ

What is Gemini Omni?

Gemini Omni is Google's multimodal video model family. It generates and edits video with audio from text, images, clips, and audio references across supported surfaces; the stable Gemini API currently supports text, image, and video inputs, while uploaded audio references are unsupported.

Is Gemini Omni available through an API?

Yes. Gemini Omni 1.1 Flash is available as the stable gemini-omni-1.1-flash model in the Gemini API. The earlier gemini-omni-flash-preview model is scheduled to shut down on September 30, 2026.

What resolutions does Gemini Omni support?

The current model supports 360p, 720p, 1080p, and 4K options. Google describes 1080p and 4K as upscaled output; 360p is intended for faster, lower-cost previews.

How long are Gemini Omni videos?

The stable model page documents 3–10 second generated outputs at 24 FPS. Scene extension can add 10-second segments up to a 40-second sequence.

How much does the Gemini Omni API cost?

Google lists $1.50 per 1M input tokens, $9 per 1M text-output tokens, and $17.50 per 1M video-output tokens. At the published 720p token rate, video output is approximately $0.10 per second.

Are Gemini Omni videos watermarked?

Yes. Generated clips carry Google's invisible SynthID watermark for AI provenance.

Is Gemini Omni the same as Gemini Flash?

No. Gemini Flash models such as 3.7 Flash primarily return text for reasoning, coding, and agent tasks. Gemini Omni Flash produces video with audio and is designed for creative generation and editing.

Sources

Version History

Gemini Omni 1.1 Flash

Current Version

Released on August 27, 2026

View Update
+What's new
3 updates
  • Move production integrations to stable gemini-omni-1.1-flash before the original gemini-omni-flash-preview model shuts down on September 30, 2026
  • Extend scenes in 10-second increments up to 40 seconds, interpolate between first and last frames, and use video references up to three seconds for tighter creative control
  • Preview at 360p up to 60% faster and at one-third the 720p cost, or deliver 720p, upscaled 1080p, and upscaled 4K output at about $0.10 per second for 720p API video

Gemini Omni Flash (API Public Preview)

Released on June 30, 2026

+What's new
2 updates
  • Generate 3–10 second 720p videos from text or still images through gemini-omni-flash-preview, then edit uploaded or generated clips with natural-language prompts
  • Iterate on generated clips through conversational editing before the stable 1.1 API added scene extension, frame interpolation, and broader resolution controls

Gemini Omni Flash (Consumer Launch)

Released on May 19, 2026

+What's new
2 updates
  • Create and edit video conversationally from text, image, video, and supported audio references in Gemini app and Google Flow
  • Reach creators through YouTube Shorts and YouTube Create while Google prepared developer and enterprise API access for the following weeks

Track Gemini Omni in ToolWorthy Weekly

Important tool updates, better alternatives, and selected AI signals in one weekly brief.

Weekly only. Unsubscribe anytime.