Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

LTX-2.3: Image to Video With Audio

Turn a still into a 1080p clip with synchronized sound using LTX-2.3, the open-weight 22B model from Lightricks. Upload an image, write the shot, hit run.

109

Generates in about 4 mins 38 secs

Nodes & Models

LoadImage
GetNode
Note
ClownOptions_SDE_Beta
SaveVideo
LatentUpscaleModelLoader
PrimitiveInt
RandomNoise
Fast Groups Bypasser (rgthree)
KSamplerSelect
SharkOptions_Beta
ManualSigmas
LTXVAudioVAELoader
PrimitiveBoolean
LTXAVTextEncoderLoader
CheckpointLoaderSimple
CLIPTextEncode
ClownSampler_Beta
ComfyMathExpression
SetNode
LoraLoaderModelOnly
easy showAnything
LTXVConditioning
LTXVEmptyLatentAudio
LTX2LoraLoaderAdvanced
CFGGuider
EmptyLTXVLatentVideo
LTX2_NAG
LTXVConcatAVLatent
SamplerCustomAdvanced
LTXVCropGuides
LTXVLatentUpsampler
LTXVSeparateAVLatent
ImageResizeKJv2
LTXVImgToVideoInplace
PreviewImage
LTXVPreprocess
LatentUpscaleBy
LatentInterpolate
CreateVideo
LTXVAudioVAEDecode
ResizeImageMaskNode
VAEDecodeTiled

ABOUT THE WORKFLOW

Animate a Still With Sound Upload an image and write the shot you want, including the sound and the score. LTX-2.3 builds the picture and the audio together, then runs the result through two upscaling passes and returns a 1080p clip with sound already on the track.

Model

  • LTX-2.3 by Lightricks. An open-weight 22 billion parameter model that generates video and synchronized audio in one pass, with a separate video stream and audio stream that attend to each other while they run.

  • Distilled speed LoRA. Cuts the base pass down to eight sampling steps.

  • Detailer LoRA. Adds micro texture to faces, fabric, and edges on the finishing pass.

  • Two spatial upscalers. Take the render from a fast base pass up to full 1080p in stages.


HOW IT WORKS

Step 1. Upload your image The frame the clip starts from. It sets the subject, the scene, and the framing. Works great with: landscapes · characters · products · concept art

Step 2. Write the shot The prompt is split into sections. One describes what happens on screen, one describes the soundscape, and one describes the music. Fill in all three and the model treats them as one brief.

Step 3. Set your output size and length Width and height set the finished clip. Length counts frames, so 240 at 24 frames per second gives you 10 seconds.

Step 4. Run it from text instead (optional) Flip the T2V switch to true and the image stops driving the render. The model builds the whole clip from your prompt.

Step 5. Hit run and download The clip renders small, gets upscaled twice, and comes back as a 1080p video file with sound, saved under video/LTX_2.3_i2v. Ready for: Premiere · DaVinci Resolve · CapCut · After Effects

First time? Leave every setting as-is. The defaults (1920 x 1080 · 24 fps · 240 frames) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard clip (most people) — 1920 x 1080 · 24 fps · 240 frames. The right starting point for almost everyone.

  • Faster test runs — Halve the width and height. Every pass in the chain is derived from those two numbers, so dropping them cuts the base render, both upscales, and the finishing pass at once.

  • Shorter or longer clip — Change the length value. It counts frames rather than seconds, so divide your target by the frame rate. 120 frames at 24 fps gives you 5 seconds.

  • Every run returns the same clip — The noise seeds ship on fixed. Change a seed number, or switch it to randomize, when you want a different take on the same prompt.

  • Sound is not landing — The soundscape and the music live in their own sections of the prompt. Rewrite those lines directly rather than burying the audio in the visual description.

  • Texture looks soft — The detail LoRA sits at 0.75. Raise it toward 1 for more micro texture in fabric, hair, and skin. Lower it if edges start to look over-sharpened.

  • The clip drifts off the source image — Say what has to stay the same. Naming the geography, the lighting, and the character design holds the model to your still across the full length.

  • Want text to video — Set the T2V switch to true. Your uploaded image stops conditioning the render.

Prompt: Write it in three parts. Describe the action and the camera in the visual section, keeping the scene locked with lines like "no cuts, no change of location, consistent lighting." Put diegetic sound in the soundscape section, like wind and waves. Put the score in the music section, like "epic orchestral, deep drums, building." A prompt that names all three lands closer than one long paragraph.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Cinematic Shots Take a concept frame and get a 10 second camera move with wind, ambience, and a score already mixed in.

📢 Ads and Product Films Turn a product still into a finished spot at 1080p without a shoot or a separate sound pass.

🎞️ Previz and Animatics Block out how a shot moves and sounds before it goes into a real production schedule.

🎵 Sound-Led Scenes Write the soundscape and the music as part of the brief when the audio matters as much as the picture.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clear stills with a readable subject and depth

  • Prompts that name the camera move and its speed

  • Explicit continuity lines like "no cuts, no change of location"

  • Soundscape and music written as their own sections

⚠️ May produce softer results

  • Prompts that describe the still instead of what happens next

  • Several unrelated actions crammed into one shot

  • Rapid motion sustained across the full length of the clip

  • Leaving the audio sections empty and hoping for sound


FAQ

What is LTX-2.3? LTX-2.3 is an open-weight video model from Lightricks, built as a 22 billion parameter diffusion transformer with a video stream and an audio stream that cross-attend while they generate. It produces synchronized picture and sound in a single pass and supports text to video, image to video, and keyframe work. It is the successor to LTX-2, with a rebuilt VAE, better prompt adherence, and new spatial upscalers.

Does LTX-2.3 generate audio with the video? Yes. Sound is generated in the same pass as the picture rather than added afterwards, which is why the prompt has its own sections for the soundscape and the score. Ambience, effects, and music come back on the track of the finished file.

Is LTX-2.3 free for commercial use? The weights and code are public, and Lightricks makes them free for commercial use for companies under 10 million USD in annual revenue. Above that threshold a commercial license from Lightricks is required. You will see LTX-2.3 described as Apache 2.0 in places, which does not match Lightricks' own licensing page, so check the license text before you build on it.

Can this workflow do text to video as well? Yes. There is a switch built into the workflow that turns off image conditioning. Set it to true and the model generates from your prompt alone, with no uploaded still involved.

Why does the clip render small before it comes out at 1080p? The workflow generates at a fraction of your target size, then runs two upscaling passes with dedicated LTX-2.3 upscaler models, resampling and refining as it climbs. That is faster than sampling at full resolution from the start and it keeps detail steadier through motion.

What resolution and length does this workflow produce? It ships at 1920 x 1080, 24 frames per second, 240 frames, which comes to a 10 second clip. Width, height, frame rate, and length are all editable, and the render chain scales off the width and height you set rather than fixed numbers.

How to run LTX-2.3 online? You can run LTX-2.3 online through Floyo. No installation, no setup, no 22B model to download. Open the workflow in your browser, upload an image, write the shot, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.

Read more

N