Back to Blog

AI

Gemini Omni Flash 1.1 Is Live in MITPO: Turn a Prompt or Photo Into Video With Sound

Gemini Omni Flash 1.1 turns a prompt, photo, or clip into real video with synchronized sound - one model, not a pipeline. See what's new and how marketers use it in MITPO.

August 28, 202612 min read
Gemini Omni Flash 1.1 Is Live in MITPO: Turn a Prompt or Photo Into Video With Sound hero image

Gemini Omni Flash 1.1: The One-Model Video Studio for Marketers

For two years, "making an AI video" meant running a relay race. Generate a silent clip in one tool. Write a script in a doc. Synthesize a voiceover in a third app. Drop it all into an editor to sync and stitch. Every hand-off lost time, context, and a little more of the original idea.

Google's Gemini Omni Flash 1.1 collapses that relay into a single runner. Text, a photo, or an existing clip goes in - and a finished video with synchronized sound comes out of the same model. It reached general availability on August 27, 2026 (model id gemini-omni-1.1-flash), and MITPO migrated to it the same day. It is live now in Creative Studio.

This guide covers what Omni Flash actually does, the three things that make it different, every generation mode with the marketing job it solves, and an honest read on what we serve today versus what is still on the roadmap.


Three things that make Omni Flash different

Plenty of models turn text into video. Omni Flash's model card calls out three capabilities that separate it from the previous generation - and all three matter more for marketing than for a tech demo.

1. Native multimodality

It processes text, image, audio, and video simultaneously in one model rather than bolting an audio model onto a video model. That is why the output feels cohesive: the sound is generated with the picture, not laid over it afterward. For a marketer, it means a product photo can become a clip that already has ambience, footsteps, or a spoken line - no separate voiceover pass.

2. Conversational editing

Because Omni Flash runs on Google's Interactions API, you refine a video the way you would brief a junior editor: in plain language, one change at a time. "Make the lighting warmer." "Swap the background to a kitchen." The model applies the change while preserving everything you did not mention, and it remembers the video's state across turns - so you are not re-rendering from scratch or re-describing the whole scene every time.

3. World knowledge

Omni Flash pairs an understanding of physics with Gemini's grasp of history, science, and culture. In practice that is the difference between a clip that merely looks real and one that behaves plausibly - liquids pour, fabric settles, a product sits on a surface the way a viewer expects. It bridges photorealism and storytelling, which is exactly where most marketing video lives.

Text, a photo, and a clip all flowing into one model that outputs a single video with sound


Every mode, and the marketing job it does

Omni Flash is one model that picks the right behavior from your inputs (you can force it with a task parameter, but Google recommends letting the prompt decide). Here is the full set, mapped to the job a marketer would hire it for.

Text to video → the "we need a clip by end of day" job

Describe a scene - subject, camera movement, lighting, mood - and get a video with audio back. The more specific the prompt (camera push-in, time of day, sound design), the better the result. This is your fastest path from a campaign idea to a usable social clip.

Image to video → the "bring the product shot to life" job

Drop a photo - a product shot, an illustration, a hero image - and the model animates it. Google's own guidance is blunt about the difference between a good and bad result: "make it move" produces weak output, while a detailed description of the camera movement, subject motion, and environmental effects produces a compelling one. Treat the photo as the anchor and the prompt as the direction.

First-and-last-frame interpolation → the "pin the shot" job

Provide a starting image and an ending image, describe the transition, and Omni Flash animates the scene from the first frame to the last. This is real directability, not a lucky seed - a marketer can lock the opening and closing frame of an ad and let the model handle the motion between them.

Subject reference → the "keep it on-brand" job

Hand the model reference images of a specific subject - a mascot, a product, a spokesperson stand-in - and it keeps that subject consistent across the generated video. Consistency is the single hardest thing in AI video and the thing brand teams care about most.

Conversational edit → the "iterate without starting over" job

Generate a video, then keep talking to it. Adjust lighting, swap a background, change a detail - each turn builds on the last using a previous_interaction_id, so the model preserves the parts you like. This turns AI video from a slot-machine ("regenerate and pray") into an actual editing session.

A video refined across iterations, one change applied while the rest of the frame is preserved

Extend → the "make it longer" job

Append a seamless 3-10 second continuation to a clip - either one the model generated or one you uploaded. "Continue the scene: the camera pans across the mountains." Extension appends to the end of a clip only, and uploaded inputs are capped at 10 seconds, but for stretching a short hook into a full spot it removes an entire re-render.


Getting better results: prompt tips

Omni Flash rewards specificity. Google's own guidance and our early runs point to the same handful of habits:

  • Describe the motion, not just the scene. "A watch on a table" is a still. "Slow 3-second push-in on a watch, morning light raking across the dial, faint ticking" is a shot. Name the camera movement, the subject motion, and the environmental effects.
  • Set lighting and mood explicitly. Time of day, warm vs. cool, soft vs. hard light - these do most of the work in making a clip feel like your brand rather than a stock render.
  • For photo-to-video, treat the image as the anchor and the prompt as the director. Tell it what should move and how; "make it move" is the weakest prompt you can give it.
  • Iterate in small steps. With conversational editing you get a better result by changing one thing at a time - lighting, then background, then pacing - than by rewriting the whole brief at once.
  • Lean on prompting before the task field. Setting the mode explicitly adds constraints; let the prompt choose the behavior unless you need to force it.

The parameters that matter

  • Aspect ratio: 16:9 (landscape, default) or 9:16 (portrait) - the two that map to a landing page and a vertical feed.
  • Resolution: 360p, 720p (default), 1080p (upscaled), and 4K (upscaled).
  • Inputs: up to 10 images or 3 short videos per prompt.
  • Duration: clips generate in the 3-10 second range - the exact length social and ad placements consume.

A note on the resolutions: 720p is the native tier and the one most placements use. 1080p and 4K are upscaled from that base, so pay for them when the deliverable genuinely needs the pixels, not by default.


What this changes for a marketing team

An AI marketing operating system earns its keep by removing hand-offs, and Omni Flash removes several at once:

  1. No separate voiceover pass - the clip arrives with sound.
  2. No "make it move" tool switch - a static product photo becomes a moving ad from the same surface you wrote the copy in.
  3. No storyboard-to-render gap - first-and-last-frame control lets you pin the shots and let the model fill the motion.
  4. No regenerate-and-pray loop - conversational editing means you refine, not reroll.

For a small team, that is the difference between "we should make a video for this campaign" and "the video is already in the review queue."


How to use Omni Flash 1.1 in MITPO today

Inside Creative Studio:

  1. Open the video surface and pick Omni Flash.
  2. Type a prompt, drop a photo to animate, or add a clip to transform.
  3. Generate. You get a 720p video with sound, ready to schedule into your channels.

We serve 720p with audio today, because that is what social and ad placements actually run. The model's higher tiers (1080p, 4K) and clip extension are on our roadmap - and we are wiring the pricing tier-by-tier first, so you never get charged for a resolution you did not ask for.


A few honest caveats

No model is magic, and we will not pretend Omni Flash tops every leaderboard - dedicated video models still trade the crown back and forth on raw fidelity. A few real limits worth knowing:

  • Editing or extending uploaded videos is restricted in the EEA, Switzerland, and the UK (videos the model generated can still be edited there).
  • Uploaded video inputs for editing or extension are capped at 10 seconds.
  • Uploading or editing media with minors or certain recognizable people is not supported.

What Omni Flash is genuinely great at is being one model that turns inputs a marketing team already owns - a product photo, a line of copy, a rough clip - into a finished, sound-on video without a tool-chain in between. Inside a connected marketing OS, that integration is worth more than a single benchmark point.

Want to see it on your own product photo? Open Creative Studio and drop one in.

Share this article