← All posts / Models

Gemini Omni 1.1 Flash Ships to Production: 40-Second Scenes, Keyframe Control, and 4K Output

Google DeepMind's Gemini Omni 1.1 Flash brings studio-grade generative video to the Gemini API — 10 seconds of scene-extension context, first/last-frame interpolation, 360p drafts at a third of the cost, and 4K upscaling.

Gemini Omni 1.1 Flash Ships to Production: 40-Second Scenes, Keyframe Control, and 4K Output

Three months after Gemini Omni first turned heads at Google I/O as the model that “creates anything from any input,” Google DeepMind has shipped the update that moves it from demo to production tool. On August 27, 2026, the team released Gemini Omni 1.1 Flash, a generally available update to the Gemini API that adds the controls professional video workflows actually demand: longer coherent scenes, precise keyframe direction, cheap iteration drafts, and 4K finishing.

For developers building generative video tools, this release is less about raw capability and more about controllability — the difference between a model that produces impressive clips and one that can be directed.

What’s new in Omni 1.1

Scene extension with real context

The headline feature is scene extension. Omni 1.1 can now analyze up to 10 seconds of prior context when continuing a video — a significant leap from earlier versions, which referenced only the final second of footage. In practice, this means extended footage holds visual consistency and narrative adherence far better: characters keep their faces, lighting stays continuous, and the story doesn’t drift when you branch in a new creative direction.

Videos can be extended in 10-second increments up to a cumulative total of 40 seconds. Google’s own examples show the range this enables — from a dolly-zoom that stretches a corridor of stone pillars while keeping a character’s face locked at constant size, to dialogue continuations where a new character walks into an ongoing scene and speaks.

Notably, the API exposes this through an interactions paradigm: you pass a previous_interaction_id and ask the model to continue the scene, optionally specifying output resolution. That’s a conversation-style interface for video, which should feel natural to anyone who has built on chat APIs.

First and last frame interpolation

Omni 1.1 lets you specify the starting and ending frames of a shot, and generates continuous video between them. The feature targets the hardest parts of editing: complex camera orbits, zoom transitions, whip-pans between subjects, and seamless loops. Google’s examples include a whip-pan from a drummer to a saxophonist and a ballet dancer, executed as “one continuous shot, no jump cuts” — the kind of instruction that used to be a wishlist item for video models.

360p drafts: iterate 60% faster at a third of the cost

Creative work is iterative, and rendering every experiment at full resolution is wasteful. Omni 1.1 supports lightweight 360p previews that generate up to 60% faster and at roughly one-third the cost of standard 720p output (based on Google’s system throughput measurements). The intended workflow is obvious and sensible: storyboard and explore in 360p, then re-render the winning take at high resolution.

4K output

When a concept is locked, Omni 1.1 can upscale to 1080p or 4K for production-ready deliverables. Combined with the draft workflow, Google is explicitly targeting the full pipeline — prototype, iterate, finish — inside one model.

Video references in multimodal input

The model now accepts up to three seconds of reference video alongside images and text in a single prompt. That enables character and motion consistency use cases that were previously painful: Google’s demo applies the dance moves from three reference videos of human dancers onto a dog, an octopus, and a bear, composited into one continuous shot. For anyone building avatar or brand-character tools, reference-driven consistency is the feature that makes AI video usable in production.

Pricing and availability

Omni 1.1 Flash (model ID gemini-omni-1.1-flash) is available now through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Pricing on the Agent Platform is set at $1.50 per million input units (text, image, video, audio), $9.00 for text output (response and reasoning), and $17.50 for video output. Existing users of the older gemini-omni-flash-preview endpoint should note its deprecation on September 30, 2026 — migration to 1.1 is straightforward since the API shape carries over.

Who’s already shipping with it

Google lined up a credible customer bench for the launch. Adobe has integrated Gemini Omni Flash into Adobe Firefly for video editing. Figma Weave uses it as one of the strongest video models on its canvas, with its creative director noting that extensions, richer references, and 4K take teams “beyond generating videos to truly directing them.” Runway offers Omni Flash alongside its own models, describing it as fitting naturally into prompt-image-video workflows. GMI Cloud highlighted the model’s accuracy for educational and explanatory content, where details must hold up under scrutiny.

Why this matters

The generative video market has been in a strange state for the past year: models produce increasingly stunning seconds, but professional pipelines still treat AI output as raw material requiring heavy cleanup. The gap hasn’t been raw quality — it’s been directability. Directors and editors think in terms of keyframes, scene continuity, and iteration cost, and most video models have ignored all three.

Omni 1.1’s feature list reads like a direct answer to that critique. Ten seconds of extension context addresses continuity. First/last-frame control addresses keyframe thinking. 360p drafts address iteration economics. 4K upscaling addresses finishing. Each feature alone would be an incremental update; together they sketch a credible production workflow rather than a novelty demo.

The competitive context sharpens the point. OpenAI’s Sora line, Runway’s Gen series, and Kling have all pushed quality; the differentiation battle is moving to workflow integration — which is exactly where Google’s distribution advantage (AI Studio, Vertex-style enterprise tooling, plus partners like Adobe and Figma) compounds. When a model is one API call inside tools creative teams already use, switching costs favor incumbency.

There are open questions worth watching. Forty seconds of cumulative scene length is a big step but still short of narrative formats; consistency across many minutes remains unsolved. Pricing at $17.50 for video output means heavy iteration at full resolution adds up fast — making the 360p tier not a convenience but an economic necessity. And the September deprecation of the preview endpoint signals Google intends to keep a fast cadence, which is good for capability but demands maintenance from API consumers.

Still, the trajectory is clear. With Omni 1.1 Flash, generative video crosses from “impressive” to “directable” — and that’s the threshold that matters for real production work.