Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

LTX-2.3 · Face Consistent Image to Video

Animate a face and keep identity locked through the clip with LTX-2.3 and a face LoRA. Upload a portrait, describe the scene, get 1080p video with audio.

9.1k

Gen time: ~2 min 51 secs

Nodes & Models

LoadImage
LatentUpscaleModelLoader
KSamplerSelect
LTXVAudioVAELoader
ManualSigmas
PrimitiveFloat
CheckpointLoaderSimple
RandomNoise
LTXAVTextEncoderLoader
INTConstant
LoraLoaderModelOnly
PrimitiveBoolean
LTXVSeparateAVLatent
CLIPTextEncode
LTXVConditioning
CFGGuider
LTXVConcatAVLatent
LTXVEmptyLatentAudio
EmptyLTXVLatentVideo
LTX2_NAG
VAEDecodeTiled
PreviewImage
SamplerCustomAdvanced
LTXVLatentUpsampler
LTXVImgToVideoConditionOnly
LTXVAudioVAEDecode
ImageResizeKJv2
LTXVPreprocess
easy cleanGpuUsed
LTXFloatToInt
easy showAnything
GetImageSize
VAEDecode
CR Prompt Text
VHS_VideoCombine
MathExpression|pysssss

ABOUT THE WORKFLOW

Animate a Face Without Losing Identity Upload a clear face photo and describe the scene. A face consistency LoRA locks the identity while LTX-2.3 animates it, generates audio in the same pass, upscales to 1080p, and saves two versions: a fast preview and a tiled HD decode.

Model

  • LTX-2.3 22B by Lightricks. The open-weight 22 billion parameter model with a Gemma 3 12B text encoder that generates picture and audio together.

  • LiCon VBVR face consistency LoRA. A community adapter that pins facial identity across frames so the face stays recognisable through motion.

  • Distilled speed LoRA. Cuts the base pass down from the full schedule.

  • Spatial x2 upscaler. Takes the render from the draft resolution up to the final 1080p output.


HOW IT WORKS

Step 1. Upload your reference image A clear face photo. Front-facing with good lighting gives the strongest identity lock. Works great with: headshots · character portraits · actor photos · ID-style shots

Step 2. Write the scene Describe the action, the camera, and the setting. Detailed prompts work best with this model.

Step 3. Hit run The model generates 121 frames with the face locked, upscales them, decodes audio, and saves two clips: a fast preview and a tiled HD version.

Step 4. Download Both clips save under Ltx23 with a date-stamped filename. The HD version includes audio. Ready for: Premiere · DaVinci Resolve · CapCut · After Effects

First time? Leave every setting as-is. The defaults (1920 x 1080 · 8 seconds · 30 fps · fixed seed) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard face clip (most people) — 1920 x 1080 · 8 seconds · 30 fps · fixed seed. The right starting point for almost everyone.

  • The face drifts during motion — Use a cleaner, front-facing reference. Side profiles and extreme angles weaken the identity lock.

  • Want a shorter or longer clip — Change the duration. At 30 fps, 8 seconds produces 240 frames. Shorter clips finish faster and use less memory.

  • Want text-to-video instead — Set the Bypass I2V switch to true. The model generates from the prompt alone without image conditioning.

  • Want a different take — Change the seed number. Both seeds ship on fixed, so the same image, prompt, and seeds return the same clip.

  • The HD decode is slow — That is the tiled VAE decode running at full 1080p. The fast preview saves alongside it for quick checks. Use the fast version for iteration and the HD version for delivery.

  • Camera moves too fast — Name the speed in the prompt. "Slow push in" or "gentle orbit left" gives the model a pacing target.

Prompt: Describe the scene around the face, not the face itself. The LoRA handles identity. "She stands in a sunlit garden, turns slowly to camera, wind moves her hair, soft ambient light, camera gently pushes in" gives you more than "a woman turns her head." Name the camera move, the lighting, and the mood.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎭 Character Animation Bring a character portrait to life while keeping their face locked through every frame.

📢 Spokesperson and Presenter Clips Animate a headshot into a short video where the face stays consistent for ads, social, or internal comms.

🎬 Actor Previsualization Test how a cast photo moves in a scene before committing to a shoot or a full CG render.

📸 Social Content From One Photo Turn a single portrait into a short, polished video with sound for Reels, TikTok, or Shorts.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Front-facing, well-lit face photos

  • Prompts that describe the scene and action rather than the face

  • Moderate camera moves like slow push-in or gentle orbit

  • Clips of 4 to 8 seconds

⚠️ May produce softer results

  • Side profiles or extreme angles on the reference image

  • Fast camera moves that turn the face away from camera

  • Readable text in the scene

  • Reference images with heavy occlusion or accessories covering the face


FAQ

What is the face consistency LoRA? LiCon VBVR is a community-trained adapter that pins facial identity across frames during video generation. It reads the reference face at the start and enforces that identity through every frame the model generates, so the face stays recognisable even as the subject moves and the camera changes angle.

Does this workflow generate audio? Yes. LTX-2.3 generates picture and audio in the same pass. The HD version includes the audio on its track. Describe the sound in the prompt if you want specific effects or ambience.

Why does the workflow save two videos? The fast preview decodes at a lower tile size for speed. The HD version uses tiled VAE decoding at full 1080p, which takes longer but produces a sharper result. Use the fast version for iteration and the HD version for delivery.

Can I use this as text-to-video instead? Yes. Set the Bypass I2V switch to true. The model runs from the prompt alone without image conditioning, and the face LoRA still applies to faces the model generates.

What resolution and length does this produce? 1920 x 1080 at 30 fps, 8 seconds by default. The model generates 121 frames at draft resolution, upscales them with a spatial x2 model, and the final output runs at the width and height you set.

Is this free for commercial use? The LTX-2.3 model is released under the LTX-2 Community License, which is free for organizations under 10 million USD in annual revenue. Above that threshold a commercial license from Lightricks is required. The face LoRA and distilled LoRA are community adapters with their own terms. Check each component before building on it.

How to run LTX-2.3 face consistent video online? You can run LTX-2.3 face consistent video online through Floyo. No installation, no setup, no model downloads. Open the workflow in your browser, upload a face photo, write the scene, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Upload a face photo, describe the scene, and run it. The identity stays locked.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N