Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

LTX 2.3 Two-Pass · Image to Video For Anime

Upload an image and LTX Video 2.3 22B generates a 5-second video at 1080p with synchronized audio, using a two-pass pipeline that renders at low resolution first then upscales in latent space for sharp, detailed output.

103

Gen time: ~3 min 14 secs

Nodes & Models

PrimitiveBoolean
GetNode
LoadImage
SaveVideo
PrimitiveInt
RandomNoise
LTXAVTextEncoderLoader
LTXVAudioVAELoader
LatentUpscaleModelLoader
ManualSigmas
KSamplerSelect
CheckpointLoaderSimple
LTXVConcatAVLatent
CFGGuider
SamplerCustomAdvanced
LoraLoaderModelOnly
LTXVPreprocess
ComfyMathExpression
LTXVAudioVAEDecode
CLIPTextEncode
LTXVEmptyLatentAudio
LTXVSeparateAVLatent
CreateVideo
ImageResizeKJv2
VAEDecodeTiled
LTXVConditioning
EmptyLTXVLatentVideo
LTXVImgToVideoInplace
LTXVCropGuides
LTXVLatentUpsampler
SetNode
FloyoStickyNote
ResizeImagesByLongerEdge

ABOUT THE WORKFLOW

Generate Video with Two-Pass Upscaling
Upload a starting image and describe the scene. Pass 1 generates base video at 768x512. Pass 2 upscales it to 1920x1080 using a dedicated spatial upscaler with a refinement sampling step. Audio is generated alongside the video. The output is a 5-second 1080p MP4 at 24fps with sound. A toggle switches between image-to-video and text-to-video mode.

Model

  • LTX Video 2.3 22B Dev by Lightricks. An open-source audio-visual video model that generates synchronized video and audio in a single architecture, paired with a Gemma 3 12B text encoder and a distilled LoRA for fast generation.

  • LTX 2.3 Spatial Upscaler x2. A dedicated latent upscaler that doubles resolution between passes while preserving temporal consistency and fine detail.


HOW IT WORKS

Step 1. Upload your image
The image that becomes the first frame. The workflow preprocesses and resizes it automatically.
Works great with: cinematic stills · portraits · concept art · landscapes · product shots

Step 2. Write your prompt
Describe the motion, camera work, and atmosphere across the full 5-second duration. "She walks steadily forward, head held level, gaze fixed ahead. The camera performs a single smooth push-in from a wider shot" gives the model clear direction.

Step 3. Hit run
Pass 1 generates at 768x512 (97 frames). The spatial upscaler doubles the latent resolution. Pass 2 refines at full 1080p. Audio generates and decodes separately. The full pipeline produces a 24fps MP4 with sound.

Step 4. Download
The output is a 1920x1080 MP4 at 24fps with synchronized audio.
Ready for: Premiere · DaVinci Resolve · After Effects · broadcast · social media

First time? Upload an image, write a motion prompt, and hit run. Leave all settings as-is. The two-pass pipeline is preconfigured.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard 1080p video — 1920x1080, 121 frames (~5 seconds), 24fps, distilled LoRA at 0.5. Upload your image, write a prompt, run.

  • Text-to-video instead of image-to-video — Toggle "Mode: Text2Video" from False to True. The workflow generates entirely from the prompt with no starting image.

  • Shorter clip — Reduce frame length from 121 to 61 (~2.5 seconds). Faster generation, lower VRAM usage. The base pass drops from 97 to about 49 frames.

  • Longer clip — Increase frame length beyond 121. More frames at 24fps extends the duration. VRAM usage scales with length.

  • Different aspect ratio — Change Width and Height (default 1920x1080). 1080x1920 for portrait. The base resolution adjusts through the math expressions automatically.

  • Faster generation — Raise the distilled LoRA strength from 0.5 toward 0.8. The model runs fewer effective steps. Trade-off: some detail softness.

  • Reproduce a result — Both pass seeds default to fixed (103 and 42). Keep them locked for identical output. Change either seed for a different interpretation.

Prompt: Describe the full 5-second arc. Front-load the camera movement, then describe subject action and atmosphere. "The camera performs a single smooth push-in, she walks forward with steady gaze, desert wind moves her dress" is sequenced and specific.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Cinematic Clips for Film and Ads
Generate broadcast-quality 1080p footage from concept stills or reference photos with the detail and stability needed for trailers, commercials, and high-end ad work.

🎨 Concept Art to Motion
Animate a single concept frame with precise camera control and atmospheric detail for pitches, pre-production, and creative direction reviews.

📺 Broadcast and Production B-Roll
Generate stock footage, B-roll, and establishing shots at 1080p with synchronized audio for broadcast, web series, and commercial production.

📱 High-Quality Social Content
Produce polished vertical or landscape video from a single image for platforms that reward high production value.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Detailed, multi-sentence prompts describing motion and camera across the full duration

  • Clear, well-lit source images at 1080p or higher

  • Smooth camera movements (push-in, dolly, tracking, orbit)

  • Scenes with environmental detail and atmospheric lighting

⚠️ May produce softer results

  • Very short or vague prompts for a 5-second clip

  • Low-resolution or heavily compressed source images

  • Extremely fast or chaotic action across the full duration

  • Frame lengths that exceed available VRAM


FAQ

What is the two-pass upscaling pipeline?
Pass 1 generates video at 768x512 base resolution with 8 sampling steps. The LTX Spatial Upscaler x2 then doubles the latent resolution. Pass 2 runs 3 more sampling steps at the upscaled resolution to refine detail. The result is 1920x1080 video that is sharper than single-pass generation at the same resolution.

Does this workflow generate audio?
Yes. LTX Video 2.3 generates synchronized audio alongside the video using a dedicated Audio VAE. The audio encodes and decodes in parallel with the video pipeline. The final MP4 includes the generated soundtrack.

Can I use this for text-to-video without a starting image?
Yes. Toggle "Mode: Text2Video" from False to True. The workflow generates entirely from the prompt with no image input required.

What resolution and duration does the output have?
Final output is 1920x1080 (landscape) at 24fps with synchronized audio. Duration is approximately 5 seconds (121 frames). Both resolution and duration are adjustable through the primitive nodes.

How does this compare to the LTX 2.3 Advanced Series workflow?
Both use multi-pass generation with spatial upscaling. This workflow uses a cleaner two-pass structure with Set/Get node organization for easier customization. The Advanced Series uses three passes with additional NAG, IC LoRA Detailer, and ClownSampler. Use this workflow for reliable 1080p output. Use the Advanced Series for maximum quality with more tuning options.

Is LTX Video 2.3 licensed for commercial use?
LTX Video 2.3 is an open-source model by Lightricks. Check the current license terms on the model page for commercial use in your specific project.

How to run LTX 2.3 two-pass image to video online?
You can run LTX 2.3 two-pass image to video online through Floyo. No installation, no setup, no local GPU needed. Open the workflow in your browser, upload your image, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Upload your image, describe the scene, and hit run. The two-pass pipeline handles the rest.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N