Video generation usually treats revision as amnesia. Ask for a different camera move and the model rebuilds the scene from scratch, perhaps changing the face, room, and weather along the way. Gemini Omni 1.1 Flash treats revision more like a conversation with an editor. The previous clip remains part of the discussion.
The basic shot is still short: three to ten seconds. Version 1.1 can examine up to ten seconds of prior footage before adding another ten-second segment, and extensions can reach roughly forty seconds in total. That is better memory, not one native forty-second take. The distinction matters because every join gives continuity another chance to drift.
First- and last-frame control adds a second kind of direction. Provide the beginning and destination, and Omni generates the movement between them. That can support a camera orbit, a transition into a screen, or a loop whose final image returns to its start. Video references add motion and style clues, although current support and reliability differ across Google surfaces.
The draft workflow may save more money than any visual flourish. A creator can generate several 360p variations for about three cents per second, compare one controlled change at a time, and only promote the successful version. Google says 360p drafts can be up to 60% faster and cost one third as much as standard 720p. Delivery can then be upscaled to 1080p or 4K.
That last verb is important. Upscaling produces a higher-resolution file; it does not travel back in time and make the original scene natively generated at 4K. Fine texture, text, hands, and complex motion can still carry errors from the underlying synthesis. Marketing often prints “4K” in large type and “upscale” in small type. A useful review gives both words the same size.
The independent family evidence is excellent. Gemini Omni Flash ranks at or near the top of several Artificial Analysis text-to-video and image-to-video preference boards, with and without audio. Arena’s early row for Omni 1.1 is also strong. These are separate Elo systems on separate scales, so their numbers cannot be averaged. Artificial Analysis has not yet isolated version 1.1; its large sample describes the broader Omni Flash endpoint.
Distribution is a genuine advantage. Omni lives in Google’s Gemini and Flow products, appears in AI Studio and developer APIs, reaches enterprise tooling, and is integrated by creative partners. API pricing is unusually legible: around ten cents per second for 720p, with cheaper drafts and more expensive upscale delivery. There is no API free tier, while consumer plans and YouTube surfaces have their own credits and conditions.
Preview status makes the product less uniform than the brand suggests. Gemini API examples use gemini-omni-1.1-flash, Enterprise Agent Platform documents a preview-suffixed ID, and the older public preview remains visible in parts of the model catalog. Audio references, C2PA credentials, and video-reference behavior vary by surface. SynthID watermarking is the broadest safe statement; universal C2PA on every output is not.
Why keep Omni second? Seedance 2.5 can attempt thirty seconds in one pass and accept a much larger stack of references. For a planned commercial scene, those controls matter. Omni is the better conversational studio: try an idea, discuss the edit, extend the result, and upscale the winner. Both score 9.2 because their strengths point in different directions. The right choice depends on whether you arrive with a storyboard or discover the storyboard while talking.