Ranked #2 Video Generation — Hollywood in a Text Box
Google DeepMind

Gemini Omni 1.1 Flash

A Google video model that remembers the conversation: generate a short scene, discuss the change, extend it, set first and last frames, draft cheaply, and upscale the approved take.

Updated August 29, 2026 Video GenerationNative AudioConversational Editing
9.2out of 10
Official Website
Best for

A Google video model that remembers the conversation: generate a short scene, discuss the change, extend it, set first and last frames, draft cheaply, and upscale the approved take.

Why It Wins

The broader Gemini Omni Flash endpoint leads or nearly leads several independent video preference boards. Version 1.1 adds up to 10 seconds of scene memory for chained extensions, first/last-frame interpolation, 360p drafts, and 1080p/4K upscaling across Google's ecosystem.

Watch out

Individual generations remain 3–10 seconds; 40 seconds means chained extension, and 4K means upscale. The API is Preview, feature support differs by surface, and independent leaderboards have not yet isolated version 1.1 from the broader Omni Flash endpoint.

01

What It Actually Is

Video generation usually treats revision as amnesia. Ask for a different camera move and the model rebuilds the scene from scratch, perhaps changing the face, room, and weather along the way. Gemini Omni 1.1 Flash treats revision more like a conversation with an editor. The previous clip remains part of the discussion.

The basic shot is still short: three to ten seconds. Version 1.1 can examine up to ten seconds of prior footage before adding another ten-second segment, and extensions can reach roughly forty seconds in total. That is better memory, not one native forty-second take. The distinction matters because every join gives continuity another chance to drift.

First- and last-frame control adds a second kind of direction. Provide the beginning and destination, and Omni generates the movement between them. That can support a camera orbit, a transition into a screen, or a loop whose final image returns to its start. Video references add motion and style clues, although current support and reliability differ across Google surfaces.

The draft workflow may save more money than any visual flourish. A creator can generate several 360p variations for about three cents per second, compare one controlled change at a time, and only promote the successful version. Google says 360p drafts can be up to 60% faster and cost one third as much as standard 720p. Delivery can then be upscaled to 1080p or 4K.

That last verb is important. Upscaling produces a higher-resolution file; it does not travel back in time and make the original scene natively generated at 4K. Fine texture, text, hands, and complex motion can still carry errors from the underlying synthesis. Marketing often prints “4K” in large type and “upscale” in small type. A useful review gives both words the same size.

The independent family evidence is excellent. Gemini Omni Flash ranks at or near the top of several Artificial Analysis text-to-video and image-to-video preference boards, with and without audio. Arena’s early row for Omni 1.1 is also strong. These are separate Elo systems on separate scales, so their numbers cannot be averaged. Artificial Analysis has not yet isolated version 1.1; its large sample describes the broader Omni Flash endpoint.

Distribution is a genuine advantage. Omni lives in Google’s Gemini and Flow products, appears in AI Studio and developer APIs, reaches enterprise tooling, and is integrated by creative partners. API pricing is unusually legible: around ten cents per second for 720p, with cheaper drafts and more expensive upscale delivery. There is no API free tier, while consumer plans and YouTube surfaces have their own credits and conditions.

Preview status makes the product less uniform than the brand suggests. Gemini API examples use gemini-omni-1.1-flash, Enterprise Agent Platform documents a preview-suffixed ID, and the older public preview remains visible in parts of the model catalog. Audio references, C2PA credentials, and video-reference behavior vary by surface. SynthID watermarking is the broadest safe statement; universal C2PA on every output is not.

Why keep Omni second? Seedance 2.5 can attempt thirty seconds in one pass and accept a much larger stack of references. For a planned commercial scene, those controls matter. Omni is the better conversational studio: try an idea, discuss the edit, extend the result, and upscale the winner. Both score 9.2 because their strengths point in different directions. The right choice depends on whether you arrive with a storyboard or discover the storyboard while talking.

02

Strengths and honest limitations

Key Strengths

  • Conversation replaces rerolling: Omni can refine or edit a generated video through natural language. The useful unit is a creative session, not a single prompt followed by a slot-machine pull.
  • Extension now remembers more of the scene: Version 1.1 examines up to 10 seconds of prior footage and can append 10-second steps to a cumulative length of about 40 seconds.
  • Draft, compare, then upscale: 360p previews cost about $0.03 per second and can be generated up to 60% faster than standard 720p. Approved work can be delivered as 1080p or 4K upscale.
  • Google puts it almost everywhere: Omni appears across Gemini, Flow, Google AI Studio, the Gemini API, Enterprise Agent Platform, partner creative tools, and selected YouTube creation experiences.

Honest Limitations

  • Ten seconds is still the native shot: A forty-second result is several generated extensions sharing context, not one uninterrupted forty-second pass. Seedance 2.5 remains stronger for native long scenes.
  • Independent evidence belongs to the family, not yet 1.1: Artificial Analysis evaluates Gemini Omni Flash. Arena lists 1.1, but early samples and live Elo can move; do not claim every family score as a version-specific result.
  • Resolution language is easy to inflate: Standard generation is 720p, while 1080p and 4K are upscaled outputs. A crisp delivery file is not proof that the scene was natively synthesized at 4K.
  • Preview surfaces do not expose one identical product: Model IDs, audio references, video-reference reliability, C2PA support, and feature availability differ between Gemini API, Flow, and Enterprise Agent Platform.
03

Benchmark Snapshot

Arena text-to-video — Omni 1.1 about 1515

The August 29 live snapshot places Omni 1.1 at the top, with the broader Omni endpoint close behind. Early live Elo is promising rather than permanent.

Artificial Analysis text-to-video with audio — Omni Flash about 1237

The broader family ranks around #2 with a large sample. This is independent family evidence, not a separate 1.1 measurement.

Artificial Analysis text-to-video without audio — Omni Flash about 1324

The broader endpoint leads the checked board, supporting visual preference quality while remaining version-ambiguous.

Per-shot output — 3–10 seconds at 720p / 24 fps

This is the documented base Omni API envelope. Version 1.1 can chain 10-second extensions to about 40 seconds cumulatively.

API pricing — about $0.03/s draft; $0.10/s at 720p

Higher delivery tiers are approximately $0.15/s for 1080p and $0.30/s for 4K upscale. The API has no free tier.

04

The Verdict

Gemini Omni 1.1 Flash enters Video at #2 with a 9.2, tied in score with Seedance 2.5. Omni has the stronger case for most people: broader access, clearer prices, conversational iteration, and better cross-board preference evidence. Seedance stays #1 because this guide gives editorial priority to its 30-second single-pass canvas and much deeper reference stack for directed production. Choose Omni when the workflow is explore, discuss, revise, and upscale. Choose Seedance when the scene must be longer and tightly specified before the camera rolls.

05

Frequently Asked Questions