Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

LTX 2.3 · Image and Text to Video For Viral SNS

Animate a photo or generate video from text with synchronized audio using LTX-Video 2.3, Lightricks' 22B open-source model. Upload an image or toggle to text-only mode, describe the scene, and hit run.

50

Generates in about 1 min 31 secs

Nodes & Models

PrimitiveBoolean
CheckpointLoaderSimple
LoadImage
LTXAVTextEncoderLoader
LTXVAudioVAELoader
LatentUpscaleModelLoader
PrimitiveInt
KSamplerSelect
SaveVideo
ManualSigmas
RandomNoise
Label (rgthree)
LoraLoaderModelOnly
Reroute
CLIPTextEncode
ComfyMathExpression
LTXVConditioning
LTXVEmptyLatentAudio
EmptyLTXVLatentVideo
CFGGuider
LTXVConcatAVLatent
SamplerCustomAdvanced
LTXVCropGuides
LTXVLatentUpsampler
LTXVImgToVideoInplace
LTXVPreprocess
LTXVAudioVAEDecode
LTXVSeparateAVLatent
CreateVideo
ImageResizeKJv2
VAEDecode
easy cleanGpuUsed
easy clearCacheAll
FloyoStickyNote
ResizeImagesByLongerEdge

ABOUT THE WORKFLOW

Animate a Photo or Generate Video from Text
Upload a photo and describe the motion you want, or switch to text-only mode with a toggle. A two-pass pipeline generates at low resolution first, then upscales. The model produces synchronized audio that matches the scene automatically. Include dialogue lines and sound descriptions in the prompt for full audio-visual control. That's it.

Model

  • LTX-Video 2.3 (22B) by Lightricks. A 22B parameter DiT-based audio-video foundation model (Apache 2.0) with native audio latent support and a Gemma 3 12B text encoder. Includes the official Distilled LoRA for faster generation and a 2x spatial upscaler for high-resolution output. Clears GPU memory automatically after each run.


HOW IT WORKS

Step 1. Upload your image
The photo you want to animate. Skip this step if you toggle to text-to-video mode.
Works great with: landscapes · portraits · product shots · environments

Step 2. Describe the scene and motion
Write what moves, how the camera behaves, and what the viewer hears. The model supports detailed prompts including camera directions, dialogue lines, and audio descriptions. "Two men look forward. Camera: rack focus. Character 1 (whisper, shocked): 'Wait, do you see that?' Audio: subtle rising music" works better than "two people talking."

Step 3. Write a negative prompt (optional)
List anything to keep out of the result, like "cartoon, video game, ugly." A default negative prompt is already loaded.

Step 4. Hit run and download
The model generates the video in two passes (low-res, then upscaled), adds synchronized audio, and returns the final result. Preview it in the workflow, then download.
Ready for: Premiere Pro · DaVinci Resolve · After Effects · any editor

First time? Leave every setting as-is. The defaults (1280×720 · 169 frames · 24 fps) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard generation (most people) — 1280×720 · 169 frames · 24 fps · fixed seeds. About 7 seconds of video with audio. The right starting point for almost everyone.

  • Want a shorter clip — Lower the frame length. At 24 fps, 97 frames is about 4 seconds. Faster generation, less credit usage. Frame counts must be divisible by 8, plus 1 (so 97, 121, 145, 169).

  • Want a longer clip — Raise the frame length. 241 frames is about 10 seconds. Longer clips take more time and credits.

  • Want higher resolution — Raise width and height. 1920×1080 delivers full HD output. The upscaler handles the refinement, but generation time increases.

  • Want portrait (vertical) video — Swap width and height to 720×1280 for 9:16 output. The model supports native portrait generation.

  • Want to generate from text only — Toggle "Text to Video" to on. No image upload needed. The model creates the entire scene from your prompt.

  • Want dialogue and sound effects — Write dialogue as character lines and describe the audio in your prompt. "Character 1 (whisper): 'Do you hear that?' Audio: wind howling, distant thunder" gives the model clear audio direction.

  • Want a different take — Change both Seed Pass 1 and Seed Pass 2. Each seed pair produces a different interpretation of the same prompt.

Prompt: Write like a script. Include camera movement, subject action, dialogue, and audio direction. Structure it with labels: "camera: slow dolly in," "dialogue: Character 1: 'Hello'," "audio: birds chirping, distant traffic." The Gemma 3 12B text encoder parses structured prompts well. Vague prompts like "make it cinematic" produce vague results.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Short Film and Narrative Clips
Generate cinematic shots with dialogue, sound design, and camera direction from a single prompt. Animate a still frame into a scene with characters speaking and environmental audio baked in.

📱 Social Media Video Content
Produce short-form video for TikTok, Reels, and Shorts. Toggle to portrait mode for native 9:16 output. Generate from text alone or animate an existing photo.

🛍️ Product and E-commerce Video
Bring product photos to life with subtle motion, ambient sound, and camera movement. A slow orbit with environmental audio turns a still product shot into a richer asset.

🎨 Pre-visualization and Storyboarding
Test how a scene reads in motion before committing to a full production. Animate concept art or storyboard frames with specific camera moves and sound cues.

🎧 Audio-Visual Prototyping
Prototype scenes with matched audio in one pass. Include dialogue lines, music cues, and sound effects in the prompt to hear how the scene sounds alongside the visuals.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Structured prompts with camera, dialogue, and audio labels

  • Environmental scenes with natural motion (wind, water, crowds)

  • Moderate camera moves (dolly, pan, rack focus)

  • Photos with clear subjects and consistent lighting

⚠️ May produce softer results

  • Short or vague prompts with no camera or audio direction

  • Extreme close-ups with rapid complex motion

  • Very long clips (consistency degrades with length)

  • Blurry, heavily compressed, or low-resolution source images


FAQ

What is LTX-Video 2.3?
LTX-Video 2.3 is a 22B parameter open-source audio-video foundation model by Lightricks, released under the Apache 2.0 license. It uses an asymmetric dual-stream Diffusion Transformer with a 14B video stream and a 5B audio stream. It generates synchronized video and audio in a single pass and supports both image-to-video and text-to-video modes.

Does LTX 2.3 generate audio automatically?
Yes. The model generates synchronized audio alongside the video in one pass. Environmental sounds, dialogue, music cues, and sound effects are produced based on your prompt. Include audio descriptions like "Audio: rain on a tin roof, distant thunder" for specific results. No separate audio generation step is needed.

Can I include dialogue in the generated video?
Yes. Write character lines in your prompt using the format: Character 1 (tone): "Line." The model generates speech-like audio matched to the visual scene. Specify the delivery (whisper, shout, calm) for better results.

What is the difference between this and the LTX 2.3 Lip Sync workflow?
This workflow is the general-purpose version. It handles camera movement, scene animation, audio generation, and text-to-video in one package. The Lip Sync workflow adds a dedicated lip sync LoRA for precise mouth movement matched to an uploaded audio track. Use this workflow for general video generation. Use the Lip Sync workflow when accurate mouth synchronization to specific audio is the priority.

What does two-pass upscaling mean?
The workflow generates at low resolution first (768×512), then runs a second pass with the 2x spatial upscaler to refine the output to your target resolution. The two-pass approach produces sharper detail in faces, textures, and edges than generating at full resolution in a single pass.

Is LTX-Video 2.3 free to use commercially?
Yes. LTX-Video 2.3 is released under the Apache 2.0 license, which allows commercial use, modification, redistribution, and fine-tuning. You can use the outputs in client work, published content, and commercial products.

How to run LTX 2.3 video generation online?
You can run LTX 2.3 video generation online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload an image or toggle to text mode, describe the scene, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A filmmaker generates a shot and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Upload a photo or toggle to text mode, describe the scene, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N