Overview
Gemini Omni 1.1 Flash is the first generally available Gemini Omni model in the Gemini API. Released August 27, 2026 as gemini-omni-1.1-flash, it turns the June public preview into a stable developer release and adds controls that matter for production creative workflows: scene extension, first/last-frame interpolation, short video references, faster 360p previews, and higher-resolution delivery options.
The model still centers on short video with audio generated from text and multimodal references. Version 1.1 is less about replacing that concept and more about making it controllable and economical across the full iteration loop—from quick composition tests to an upscaled 4K deliverable.
What's New
Scene extension with more context
Omni 1.1 can extend an existing result in 10-second increments up to a 40-second sequence. Google says the model can now analyze up to ten seconds of prior context, compared with previous models that referenced only the final second. That broader temporal window should help preserve action and narrative continuity, although creators still need to check identity, objects, and camera motion at every join.
First- and last-frame interpolation
Developers can provide a starting image and an ending image, then ask the model to generate the transition. This gives art directors stronger boundary control than an open-ended prompt and is useful for product reveals, scene changes, transformation shots, and storyboard-to-video workflows.
Video references
Omni 1.1 accepts short video references of up to three seconds to guide motion, composition, or continuity. Combined with image references, this makes it easier to specify both what a subject should look like and how it should move.
Preview and delivery resolutions
The new 360p mode is intended for fast iteration. Google reports it can be up to 60% faster and cost roughly one-third as much as 720p. Final output options include 720p plus upscaled 1080p and 4K, allowing teams to separate creative selection from delivery rendering.
Performance & Cost Benchmarks
Google publishes vendor-run comparative evaluations for Gemini Omni Flash across video editing, text-to-video, fast motion, image-to-video, and reference-to-video, alongside the workflow measurements and token rates below. These evaluations are not independent tests and are not presented as a version-isolated Omni 1.1 leaderboard.
| Metric | Official figure | Practical implication |
|---|---|---|
| 360p preview speed | Up to 60% faster than 720p | Test motion and composition before a costly final render |
| 360p preview cost | About one-third of 720p | More iterations within the same budget |
| 720p output token rate | 5,792 tokens/second | Converts the published $17.50/M output rate to about $0.10/second |
| Scene extension | 10-second increments, up to 40 seconds | Supports longer sequences but still requires continuity review |
| Output frame rate | 24 FPS | Consistent baseline for generated clips |
These are vendor-published operational figures, not independent quality tests. Actual job cost includes inputs, text output, failed generations, and every conversational revision.
Compared With Previous Version
The June 30 gemini-omni-flash-preview model established the API workflow with 3–10 second 720p generation and conversational editing. Version 1.1 keeps that foundation but changes the deployment decision materially.
| Area | Omni Flash Preview | Omni 1.1 Flash |
|---|---|---|
| Lifecycle | Public preview; shutdown Sep 30, 2026 | Generally available stable model |
| Model ID | gemini-omni-flash-preview |
gemini-omni-1.1-flash |
| Resolution | 720p | 360p, 720p, upscaled 1080p, upscaled 4K |
| Extension context | Not applicable; extension was not supported by the preview API | Up to 10 seconds of prior context |
| Sequence extension | Not supported | 10-second increments up to 40 seconds |
| Boundary control | Conversational edits and references | Adds first/last-frame interpolation |
| Video references | Up-to-3-second clips were accepted by the API schema but were not correctly processed by the preview model | Up to 3 clips, each up to 3 seconds, supported in 1.1 |
The most urgent difference is lifecycle: the old preview has a fixed shutdown date, so active integrations need to migrate even if they do not immediately use the new creative controls.
Compatibility Notes
- Replace
gemini-omni-flash-previewwithgemini-omni-1.1-flashand complete migration before September 30, 2026. - Re-run prompt and reference regression tests; stable does not guarantee identical framing, motion, or edit behavior.
- The model accepts text, image, and video inputs. Uploaded videos for editing or extension are limited to 10 seconds unless extending model-generated video through multi-turn; video references support up to three clips of up to 3 seconds each. Editing or extending uploaded videos is currently unavailable in the EEA, Switzerland, and the UK.
- Generated output is 3–10 seconds per request at 24 FPS before using extension workflows.
- 1080p and 4K are upscaled outputs, so they improve delivery resolution but do not guarantee additional native scene detail.
- The API's published Gemini Omni 1.1 pricing is paid-tier only; consumer subscriptions and YouTube access follow different quotas and terms.
- All generated output includes SynthID provenance marking.
Pricing & Plans
| API component | Price |
|---|---|
| Multimodal input | $1.50 / 1M tokens |
| Text output | $9.00 / 1M tokens |
| Video output | $17.50 / 1M tokens |
| Approximate 720p video output | $0.10 / second |
There is no free Gemini API tier for this model. Use 360p for composition and motion tests, then request higher resolution only for selected clips. Budget for retries and edits rather than multiplying duration by the headline per-second figure alone.
Who Should Upgrade / Who Should Wait
Upgrade now if:
- Your application uses
gemini-omni-flash-preview; the September 30 shutdown makes migration time-sensitive. - You need longer continuous scenes, explicit start/end frames, or motion guidance from a short clip.
- Your creative workflow can save time and money by separating 360p preview from final rendering.
- You need a stable API lifecycle rather than a public-preview dependency.
Wait or evaluate longer if:
- You require native 1080p or 4K generation rather than upscaling.
- Your work depends on perfect character consistency, complex physical interactions, or accurate rendered text.
- A mandatory SynthID watermark conflicts with the distribution requirement.
- You only need occasional consumer creation and do not need the API's programmatic controls.