LTX-2.5 for Audio to Video
Generate video from an audio file using LTX 2.5, Lightricks' 22B open-source video model. Upload audio, write a prompt, add a starting image if you want, and hit run.
Audio
Audio to Video
LTX 2.5
Video
43
Nodes & Models
LoadImage
LTXVConditioning
SolidMask
ManualSigmas
UNETLoader
ltx-2.5-22b-distilled-transformer-bf16.safetensors
PrimitiveStringMultiline
CLIPTextEncode
EmptyLTXVLatentVideo
Note
VAELoader
ltx-2.5-audio-vae-bf16.safetensors
ltx-2.5-video-vae-bf16.safetensors
LTXVImgToVideoInplace
CFGGuider
PrimitiveFloat
PrimitiveInt
LTXVPreprocess
SamplerCustomAdvanced
RandomNoise
CLIPLoader
gemma4_e2b_it_bf16.safetensors
gemma4-12b-with-proj-ltx-2.5-bf16.safetensors
ResizeImageMaskNode
LTXVConcatAVLatent
KSamplerSelect
Reroute
MarkdownNote
TrimAudioDuration
ComfyNotNode
LTXVSeparateAVLatent
LatentUpscaleModelLoader
ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors
PrimitiveBoolean
SetLatentNoiseMask
LTXVLatentUpsampler
ComfySwitchNode
LoadAudio
StringContains
ComfyMathExpression
GemmaAPITextEncode
ltx-2.3-22b-dev.safetensors
PreviewAny
VAEDecode
CreateVideo
SaveVideo
ABOUT THE WORKFLOW
Generate Video From Audio
Upload an audio file and describe the visuals you want. The model generates a video with motion synced to your sound. Add a starting image to open on a specific frame. That's it.
Model
LTX 2.5 22B by Lightricks. An open-source video model that generates video with synced audio in a two-stage pipeline: 8 fast steps at low resolution, then a 2x upscale pass in 3 more steps.
HOW IT WORKS
Step 1. Upload your audio
Load an audio file. Set the start point and duration to trim the clip before generation.
Works great with: music · sound effects · voiceovers · ambient audio
Step 2. Write a prompt
Describe the visuals and any on-screen action that should follow the audio. "A drummer playing in a dim jazz club, camera slowly orbiting" is clearer than "music video."
Step 3. Add a starting image (optional)
Load an image and enable "use image input" to open the video on that still. Leave it off to let the model generate the opening frame from the prompt alone.
Step 4. Hit run and download
The model generates a video with your original trimmed audio muxed back in. One file, ready to use.
Ready for: Premiere Pro · DaVinci Resolve · After Effects · any NLE
First time? Leave every setting as-is. The defaults (5 seconds, 24 fps, no image input) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard audio-to-video (most people) — 5s duration · 24 fps · no image input · prompt enhancement off. The right starting point for almost everyone.
Want the video to open on a specific frame — Load an image and enable "use image input." The model animates outward from that still.
Need a longer clip — Increase the duration. Keep it under 10 seconds for the best quality. Longer durations increase generation time sharply.
Want richer scene descriptions without rewriting — Turn on "enhance positive prompt." The model rewrites your prompt into a more detailed version before generating.
Want to use a specific section of your audio — Set "audio start" to the timestamp (in seconds) where you want the clip to begin. The workflow trims from that point.
The motion does not match the audio — Rewrite the prompt first. Describe what moves and how it follows the sound, rather than changing settings.
Prompt: Describe the visuals and the motion you want, not the audio. "A forest canopy swaying in wind, light filtering through leaves" works better than "nature sounds." Keep it specific about what should move on screen.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎵 Music Visualizers
Upload a track and describe the visual style. Get a synced video you can use as a music visualizer, lyric video backdrop, or social media clip.
🎬 Sound Design Previsualization
Feed sound effects or dialogue into the model and generate rough visual takes that match the audio timing. Block out a scene before committing to a full shoot.
🎙️ Podcast and Voiceover Visuals
Turn a voiceover clip into an animated visual that follows the pacing of the narration. Useful for social teasers and audiogram replacements.
🎮 Game and Interactive Audio
Generate environment footage synced to ambient soundscapes or gameplay audio. Use the output as reference material, cutscenes, or background loops.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Clear, well-recorded audio with a strong rhythm or beat
Prompts that describe on-screen motion tied to the sound
Single-environment scenes (one location, one mood)
Starting images with clean lighting and clear subjects
⚠️ May produce softer results
Very quiet or ambient audio with no clear structure
Prompts that describe the audio instead of the visuals
Durations longer than 10 seconds without an extend step
Low-quality or heavily compressed audio files
FAQ
What is LTX 2.5 and how does it handle audio?
LTX 2.5 is a 22-billion-parameter open-source video model made by Lightricks. It takes audio as a direct input alongside text and images. The audio is encoded into tokens that stay locked throughout generation, so the video motion stays synced to the sound. The original trimmed audio waveform is muxed into the final video file, not decoded from latents.
How does the two-stage generation pipeline work?
The first stage generates a low-resolution video in 8 fast steps while keeping the audio tokens frozen. The second stage upscales the video 2x and refines it in 3 more steps, re-freezing the original audio tokens. The result is a higher-resolution video with tighter audio sync than a single-pass approach.
Can I use a starting image with the audio?
Yes. Load an image and enable "use image input." The model conditions the video to open on that still and animates outward from it. If you leave image input off, the model generates the opening frame from your prompt and audio alone.
What resolution and frame rate does the output use?
The default output is 960x544 at 24 fps, upscaled 2x by the second generation stage. The frame count is calculated as 1 + floor(fps × duration / 8) × 8, so the actual clip length may differ slightly from the duration you enter.
Is LTX 2.5 open source and can I use the output commercially?
LTX 2.5 has open weights on HuggingFace under the Lightricks Community License. Review the license terms for your specific use case. FP8 variants of the model run on 12 GB VRAM for local use.
What is the difference between LTX 2.5 audio-to-video and text-to-video?
This workflow takes audio as its primary input and generates video synced to that sound. A text-to-video workflow generates video from a text prompt alone, with no audio input or sync. Use this workflow when you have a specific audio clip and want the visuals to match its rhythm and timing.
How to run LTX 2.5 audio to video online?
You can run LTX 2.5 audio to video online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your audio, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A sound designer runs a workflow and likes the result. A motion artist opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it?
Upload your audio, write a prompt, and hit run. The settings are already set.
Questions? Watch the free course or check the FAQ above.
Read more




