Text to video with audio
LTX-2.3 22B turns a single written prompt into a finished video clip complete with its own synchronized audio. Instead of generating a silent clip and then hunting for sound effects or music to layer on afterward, you describe the scene you want and the model delivers moving images and matching sound together in one pass. That makes it a strong fit for filmmakers building quick previsualizations, content creators producing short social clips, motion designers exploring ideas, and any artist who wants to see a concept move and sound alive without touching a timeline.
At its core the model reads your text prompt and builds a video from it. A good prompt describes not just what appears on screen but how it feels — for example, "A cowboy walking through a dusty town at high noon, camera following from behind, cinematic depth, realistic lighting, western mood, 4K film grain." The more you describe mood, lighting, camera behavior, and texture, the closer the result lands to what you pictured. The model can also expand your prompt automatically, filling in cinematic detail so that even a shorter description produces a rich, coherent shot.
You have broad control over the shape and length of your clip. Videos can run anywhere from a brief 9-frame moment up to 481 frames, and you set the frame rate anywhere from 1 up to 60 frames per second, with 24 frames per second — the classic film cadence — as the default. The default clip is 121 frames of landscape 16:9 footage, a natural fit for widescreen and social video, but you can choose other framings to match your project. Playback speed and length work together, so a longer frame count at a lower frame rate gives you a slower, more drawn-out moment, while a higher frame rate gives smoother motion.
One of the standout creative tools is direct camera control. Rather than hoping the model interprets a phrase like "push in on the subject," you can select from a set of real camera moves: dolly in, dolly out, dolly left, dolly right, jib up, jib down, or a locked-off static shot. These give you the kind of deliberate, repeatable camera language that filmmakers rely on, and you can dial the strength of the move up or down so it feels subtle or dramatic. This makes it far easier to get intentional motion that supports your storytelling instead of random drift.
Audio is generated alongside the video by default, and you can turn it off if you only need silent footage. You also have separate creative dials for how tightly the audio and video each stick to your prompt, and how the two balance against each other — useful when you want the soundscape to feel prominent or want to keep the focus firmly on the visuals.
To help with overall coherence and fine detail, the model uses a multi-scale approach. It first sketches the video at a smaller scale, then uses that draft to guide a larger, more detailed final render. The result is stronger consistency and cleaner detail than generating at full size in a single shot. You can further tune how closely the model follows your prompt versus how much creative freedom it takes, adjust the level of refinement for a balance between speed and polish, and use a controlled variation option that can lift quality on tricky scenes.
A negative prompt gives you a way to steer the model away from looks you don't want. By default it avoids things like cartoon and video-game aesthetics, on-screen text, watermarks, logos, subtitles, slow motion, and static, flat footage — nudging results toward a more cinematic, live-action feel. You can rewrite this to fit your own aesthetic goals.
When it comes to output, you can export in several formats to suit your workflow: standard MP4, WebM, ProRes for high-end editing pipelines, or animated GIF for lightweight sharing. Quality levels range from low up to maximum, and you can bias the file toward faster writing, smaller size, or a balanced middle ground. Acceleration settings let you trade off between generation speed and fidelity depending on whether you're iterating quickly or finalizing a hero shot.
For reproducibility, you can set a seed so that the same prompt and settings produce the same result — handy when you want to make small, controlled changes rather than starting over from scratch each time. A built-in safety checker is enabled by default.
LTX-2.3 22B is best suited to short-form cinematic video: mood pieces, concept shots, social clips, animatics, and idea exploration where having synchronized sound baked in saves a whole editing step. Because clips are measured in frames and generated as self-contained moments, it shines at single evocative shots rather than long, multi-scene narratives. Getting the most from it comes down to writing descriptive, cinematic prompts, choosing an intentional camera move, and letting the multi-scale render and refinement do their work. With its combination of prompt-driven video, native audio, and precise camera control, it gives creators a fast path from a written idea to a moving, sounding scene.
A woman kneeling in darkness, illuminated by a warm, radiant beam of light emerging from her raised hand.
Describe your video scene with motion, camera angles, and mood
Model creates cinematic motion with natural physics and lighting
Download and share your production-ready video
Showcases sustained tracking motion, vehicle dynamics, and engine audio generation in a widescreen cinematic sequence built for YouTube and film.
Demonstrates atmospheric fog dynamics, lighting transitions, and immersive natural audio for Netflix-style documentary landscape content.
Highlights the model's synchronized audio-visual generation with crowd sound, dynamic stage lighting, and energetic camera motion for landscape cinema.
“Cinematic reveal of a sleek black luxury sports car in a dark studio. Camera starts close on the chrome badge, slowly pulling back while orbiting 180 degrees around the vehicle. Dramatic rim lighting gradually intensifies, highlighting the car's sculptural curves and glossy finish. Reflections dance across the body as the camera moves. Dust particles float in volumetric light beams. Final wide shot reveals the full silhouette against a gradient backdrop. 8 seconds, smooth motion, 24fps cinematic quality.”
Switch to reasoning-guided synthesis today. Be the first in your industry to deliver native 4K results at 10x the speed.

Cinematic video from references
0.4 credits

Cinematic video with native audio
1.4 credits

Fast balanced text-to-video generation
1.6 credits
![Kling Video v3 Text to Video [Standard]](https://v3b.fal.media/files/b/0a8cfc9f/dei5OqFRB9HK8AgSHwk8f_9a5eea197b3045d1be55aedb0213f6f9.jpg)
Cinematic text-to-video with audio
4.2 credits

Text to video with audio
0.3 credits

Cinematic video from references
10 credits

Film-grade video with audio
0.1 credits

Frontier 2K text-to-video generation
7.3 credits

Fast cinematic video with audio
0.1 credits