Cinematic video from references
Seedance 2 Reference to Video is ByteDance's most advanced reference-driven video model, built for creators who want to steer generation with real source material rather than text alone. Instead of describing everything from scratch, you feed the model a mix of reference files, up to 9 images, 3 videos, and 3 audio clips (with a total of 12 files across all types), and weave them together with a written prompt. You point to each asset directly in your prompt using simple tags like @Image1, @Video1, or @Audio1, so the model knows exactly which face, environment, motion, or sound you want it to honor. The result is a single continuous cinematic clip that respects your references while filling in the world around them.
The standout feature is native audio. In a single generation pass, the model produces synchronized sound alongside the picture: ambient noise, sound effects, music, and even lip-synced dialogue. There is no separate step to add sound later. That means a bustling crowd scene can arrive with overlapping shouts, camera shutter clicks, distant sirens, and an idling engine already baked in and timed to the action on screen. You can leave audio generation on to get a fully finished clip, or switch it off if you prefer to handle sound yourself.
Creative control is a major strength. You direct the camera through your prompt using cinematic language, calling for dolly zooms, tracking shots, handheld movement with natural micro-shake, POV switches, and eye-level or elevated perspectives. The model interprets these directions like a cinematographer would. It also handles realistic physics, so fluid dynamics, cloth movement, and character motion behave believably. One especially useful capability is multi-shot editing: within a single generation of up to 15 seconds, the model can create natural cuts, giving you sequences that flow between shots rather than a single locked frame.
Because it accepts image references, the model excels at consistency. Provide a character reference and it can render that same face and outfit across the shot, even when translating the subject into a different visual style. Examples show a hybrid look where stylized 3D animated characters are composited seamlessly into a fully photorealistic live-action environment, with the referenced character keeping a perfectly consistent face and the exact outfit from the source image, all while interacting believably with real light, flash bounces, streetlight highlights, and cast shadows. This makes it a strong choice for anyone who needs a recurring character, a specific product, or a fixed look to stay reliable across a scene.
On output options, you can generate at 480p for faster turnaround, 720p for a balance of speed and quality, or 1080p for high-quality results. Duration is flexible from 4 to 15 seconds, or you can let the model decide automatically based on your prompt. Aspect ratio is equally adaptable: choose 16:9 for landscape, 9:16 for vertical and social formats, 1:1 for square, 21:9 for ultrawide cinematic framing, 4:3, 3:4, or auto to let the model pick. You can also request a higher-quality encode when you want a larger, cleaner file, and fix a seed value to reproduce a result you like, though minor variation may still occur.
The reference inputs have sensible limits worth knowing. Images can be JPEG, PNG, or WebP at up to 30 MB each, with up to 9 allowed. Videos can be MP4 or MOV, with a combined duration between 2 and 15 seconds, a total under 50 MB, and each video between 480p and 720p resolution, with up to 3 allowed. Audio can be MP3 or WAV, up to 3 clips with a combined duration of no more than 15 seconds and a maximum of 15 MB each. One important rule: if you provide audio references, you must also include at least one reference image or video. And no matter how you mix them, the total number of reference files across all modalities cannot exceed 12.
Who benefits? Filmmakers and video directors get a way to previsualize or produce short cinematic sequences with real camera language and finished sound. Content creators and social media producers can turn a single character reference into consistent vertical clips. Advertisers and product marketers can drop a product image into a styled, lit environment with a controlled camera move. Animators and mixed-media artists can blend stylized characters into photoreal worlds, exactly the hybrid workflow highlighted in the examples. Because it draws on your own images, videos, and audio, it is ideal whenever brand consistency, character identity, or a specific look matters more than pure text-to-video guesswork.
A few best practices follow naturally from how it works. Write prompts like a director: describe the style, the lighting and environment, the subject, the action sequence, and the audio you want, and reference your assets explicitly with @ tags so the model knows what to preserve. Keep your video references within the duration and resolution limits so they can be used cleanly, and remember that supplying audio requires at least one visual reference to anchor it. For the most polished results, lean into detailed action descriptions and camera directions, since the model responds well to that level of specificity, as seen in the elaborate crowd and convoy example. With thoughtful references and clear direction, Seedance 2 Reference to Video delivers finished, sound-complete cinematic clips that stay faithful to the material you bring to it.
A woman kneeling in darkness, illuminated by a warm, radiant beam of light emerging from her raised hand.
Describe your video scene with motion, camera angles, and mood
Model creates cinematic motion with natural physics and lighting
Download and share your production-ready video
Highlights Seedance 2's director-level camera control with complex multi-stage camera movements, atmospheric weather simulation, and dramatic landscape-scale scene dynamics suited for widescreen travel cinematography.
Demonstrates Seedance 2's ability to handle complex scene transitions, stylized physics (shattering glass, floating debris), and dramatic lighting choreography — showcasing the cut-scene narrative capability for music video production.
Showcases Seedance 2's real-world physics engine and native audio generation with environmental sound design (crunching snow, wind, breathing) — demonstrating Netflix-quality nature documentary footage with precise animal motion and atmospheric dynamics.
“Cinematic reveal of a sleek black luxury sports car in a dark studio. Camera starts close on the chrome badge, slowly pulling back while orbiting 180 degrees around the vehicle. Dramatic rim lighting gradually intensifies, highlighting the car's sculptural curves and glossy finish. Reflections dance across the body as the camera moves. Dust particles float in volumetric light beams. Final wide shot reveals the full silhouette against a gradient backdrop. 8 seconds, smooth motion, 24fps cinematic quality.”
Switch to reasoning-guided synthesis today. Be the first in your industry to deliver native 4K results at 10x the speed.
![Kling Video v3 Text to Video [Standard]](https://v3b.fal.media/files/b/0a8cfc9f/dei5OqFRB9HK8AgSHwk8f_9a5eea197b3045d1be55aedb0213f6f9.jpg)
Cinematic text-to-video with audio
4.2 credits
Text to video with audio
0.7 credits

Film-grade video with audio
0.1 credits

Text to video with audio
0.3 credits

Cinematic video from references
0.4 credits

Frontier 2K text-to-video generation
7.3 credits

Fast balanced text-to-video generation
1.6 credits

Cinematic video with native audio
1.4 credits

Fast cinematic video with audio
0.1 credits