Multimodal references to video
Gemini Omni Flash is a multimodal video generation model that turns your creative references into short videos complete with sound. What sets it apart is its ability to combine several types of input at once — text, images, audio, and video — into a single, coherent result. Instead of relying on a single starting frame or a text prompt alone, you can feed the model a rich blend of references that together guide the subject, the motion, the visual style, and even the audio of your finished clip. This makes it a versatile tool for creators who want precise control over how their videos look and sound.
At its core, the model works by taking one or more reference images and a written prompt describing the scene you want to bring to life. For example, a simple prompt like "A cat playfully batting at a ball of yarn in a sunlit living room" paired with reference imagery produces a lively animated video. You can supply up to ten reference images, giving you the flexibility to establish characters, environments, props, or visual styles that the model will incorporate into the final video. The prompt itself supports detailed, descriptive writing — up to a generous length — so you can articulate exactly how you want the scene to unfold. You can also bind specific reference images to roles within your prompt, meaning you can tell the model which image corresponds to which subject or element in your scene, giving you a stronger hand in directing the outcome.
The model is tagged for stylized transformation and lip sync, which points to some of its most distinctive strengths. Because it accepts audio as an input, it can generate video that aligns with sound — making it well suited for creating talking characters, dialogue-driven clips, and other content where visuals and audio need to feel connected. Its stylized transform capabilities mean you can push your references beyond straightforward realism into distinctive artistic looks, driven by both your imagery and your written direction.
Gemini Omni Flash gives you a handful of intuitive creative controls. You can choose between two aspect ratios: a widescreen 16:9 format that's ideal for cinematic, landscape, and desktop viewing, or a vertical 9:16 format tailored for mobile-first content like social media stories and short-form video. You also control the length of your clip, with durations ranging from three to ten seconds and a default of eight seconds — enough to capture a moment, a gesture, or a short scene while keeping your content punchy and shareable. The model outputs 720p video, delivering clean, presentable footage suitable for a wide range of creative and professional uses.
This model is a natural fit for a broad range of creative professionals. Filmmakers and video creators can use it to prototype scenes, generate short animated sequences, or explore stylistic directions quickly. Social media creators and marketers will appreciate the vertical 9:16 format for producing eye-catching short-form content, along with the ability to include synchronized audio for engaging, voice-driven clips. Designers and visual artists can transform their still artwork and reference images into moving pieces, breathing life into static concepts. Because the model handles character-driven and lip-synced content, it's also valuable for creators building narrative shorts, animated avatars, explainer videos, or expressive character animations.
The workflow is refreshingly straightforward. You provide a prompt describing what you want to see, attach at least one reference image (and up to ten), and optionally adjust the aspect ratio and duration to suit your project. The model then generates a video file that you can download and use in your creative work. Because you can layer multiple references together — combining, for instance, a character image with a background image and a style reference — you have a great deal of expressive latitude in shaping the final result. This multimodal approach is especially powerful when you have a clear vision that spans multiple visual elements, allowing the model to blend them into one unified scene rather than forcing you to work from a single source.
When it comes to getting the best results, thoughtful prompting pays off. Descriptive, specific prompts help the model understand the subject, the action, the environment, and the mood you're after. Since the model supports binding reference images to specific roles within your prompt, taking advantage of this feature helps ensure that each of your references shows up in the video the way you intend — reducing ambiguity and giving you more predictable, controllable outcomes. Choosing the right aspect ratio for your destination platform, and selecting a duration that matches the pacing of your idea, will also help your finished clips feel intentional and polished.
A few considerations are worth keeping in mind. Clips are short by design — between three and ten seconds — so the model is best suited for concise moments, loops, and short-form storytelling rather than long-form narrative footage. The output resolution is 720p, which is well suited for web, social, and preview use. And while the model accepts audio and video as inputs to help guide the result, at minimum you'll always need a text prompt and at least one reference image to get started. Within these bounds, Gemini Omni Flash offers a flexible, reference-driven way to create expressive videos with sound, making it a compelling option for creators who want to combine multiple sources of inspiration into a single moving image.
Whether you're animating a character, transforming a piece of artwork into motion, producing a quick social clip, or exploring stylized visual ideas, Gemini Omni Flash gives you an accessible and multimodal path from reference material to finished video. Its combination of image, text, audio, and video inputs, along with lip-sync and stylization strengths, makes it a distinctive tool for bringing ideas to life quickly and with a strong degree of creative control.
Add the image that you want change
Add an optional image to guide the look, character, or environment
A woman kneeling in darkness, illuminated by a warm, radiant beam of light emerging from her raised hand.
Type a prompt - Model understands the physics, lighting, and emotional intent of your scene
Click to generate your final output and download production grade video
Demonstrates cinematic landscape animation with atmospheric motion and generated ambient nature sound for wide-format storytelling.
Showcases premium product animation combining reference imagery with dynamic lighting and sound for luxury commercial reels.
“Animate as a smooth 360-degree rotation on an invisible turntable. Rotate slowly and continuously, taking 6 seconds for full rotation. Light reflections should shift naturally across the metal case and crystal. Maintain consistent dramatic lighting throughout rotation. Add subtle sparkle on diamond indices as they catch light. Keep the background static and dark. Professional product video quality.”

Switch to reasoning-guided synthesis today. Be the first in your industry to deliver native 4K results at 10x the speed.