Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

LTX 2.3 Video Inpainting · Video to Video

Replace or add objects in a video with a single generation pass using LTX-Video 2.3 and a dedicated inpainting LoRA. Upload a video, draw a mask, describe the change, and hit run. Faster than the multi-pass version.

27

Generates in about 1 min 43 secs

Nodes & Models

Note
LTXAVTextEncoderLoader
ImageScaleBy
KSamplerSelect
LTXVConditioning
CheckpointLoaderSimple
ResizeImageMaskNode
EmptyImage
RandomNoise
GetImageSize
CLIPTextEncode
Fast Groups Bypasser (rgthree)
LTXVAudioVAELoader
LTXVPreprocess
SolidMask
LatentUpscaleModelLoader
MaskToImage
ManualSigmas
LoraLoaderModelOnly
easy cleanGpuUsed
LTXVEmptyLatentAudio
ReservedRegionFrameComposer
CFGGuider
SetLatentNoiseMask
GrowMaskWithBlur
SamplerCustomAdvanced
TrimAudioDuration
easy showAnything
BlockifyMask
LTXVSeparateAVLatent
SAM3Segment
BasicScheduler
VAEEncode
VAEDecode
LTXVAddGuideMulti
LTXVCropGuides
LTXVConcatAVLatent
ImageCompositeFromMaskBatch+
GetImageSize+
LTXVAudioVAEEncode
CM_FloatToInt
ImageCompositeFromMaskBatch+
GetImageSize+
VHS_LoadVideo
VHS_VideoCombine
VHS_VideoInfo
CM_FloatToInt

ABOUT THE WORKFLOW

Replace or Add Objects in a Video (Fast)
Upload a video, mask the region you want to change, and describe what should appear there. This is the single-pass version of the LTX 2.3 inpainting workflow. It skips the upscale refinement passes, so you get results faster and at lower cost. Use it to iterate on mask placement and prompt wording before committing to the multi-pass version for final output. That's it.

Model

  • LTX-Video 2.3 (22B) by Lightricks. A 22B parameter DiT-based audio-video foundation model (Apache 2.0) with a Gemma 3 12B text encoder. Runs with a dedicated inpainting LoRA for masked region generation and two distilled LoRAs for fewer-step inference. Audio conditioning is available as an optional advanced feature.


HOW IT WORKS

Step 1. Upload your video
The video you want to edit. The workflow loads up to 121 frames at 24 fps (about 5 seconds). Shorter clips work too.
Works great with: product shots · talking heads · street scenes · VFX plates

Step 2. Draw a mask
Paint over the region you want to replace or fill. Make it slightly larger than the target object. Everything outside the mask stays untouched.

Step 3. Describe what should fill the mask
Write a prompt that describes the full scene including the new object. "A woman walking through a park holding a red umbrella, warm afternoon sunlight" works better than "red umbrella" alone. The model needs scene context to match lighting and perspective.

Step 4. Write a negative prompt (optional)
List anything to avoid in the output. Add terms specific to your shot if needed.

Step 5. Hit run and download
The model inpaints the masked region in a single pass and returns the result. Preview it in the workflow, then download.
Ready for: Premiere Pro · DaVinci Resolve · After Effects · any editor

First time? Leave every setting as-is. The defaults (8 steps · 1 guidance · seed 42) are the right starting point. Focus on the mask and the prompt.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard inpaint (most people) — 8 steps · 1 guidance · fixed seed. The right starting point for almost everyone. Focus on drawing an accurate mask and writing a detailed prompt.

  • Want higher quality output — Use the multi-pass version of this workflow instead. It adds two upscale refinement passes after inpainting, producing sharper detail in the filled region. Use this 1-pass version for iteration and the multi-pass version for final delivery.

  • The new object does not match the scene — Write a fuller prompt. Describe the entire scene, not only the new object. Include lighting, camera angle, and surrounding context.

  • The edges of the inpaint look harsh — Make the mask larger than the object you are replacing. The workflow grows and blurs the mask automatically (10px expand, 32px block size), but a bigger starting mask gives the model more room to blend.

  • Another object in the scene is interfering — Be specific about which subject is new. If there is already a person in the frame and you are adding a different person, describe both clearly so the model does not merge them.

  • Want a different take — Change the seed. Each seed produces a different interpretation of the same mask and prompt.

  • Want audio in the output — Audio conditioning is available but bypassed by default. Enable it inside the workflow to generate scene-matched audio alongside the inpainted video.

Prompt: Always describe the full scene, not only the masked region. The prompt is the most important setting in this workflow. A vague prompt produces unpredictable fills. Include the new object, the surrounding environment, lighting, and camera angle for consistent, scene-matched results.


LEARN

📹 Videos

✨ Quick links


USE CASES

🔁 Fast Iteration on Inpaint Placement
Test mask positions, prompt wording, and seed variations before committing to the multi-pass version. Run multiple takes in the time it takes the 3-pass workflow to produce one.

🎬 VFX and Object Replacement
Remove or replace objects in a shot. Mask a logo, a prop, or an unwanted element and describe what should fill the space. The model matches lighting and perspective to the surrounding footage.

🛍️ Product Placement Previews
Insert a product into existing footage for quick review. Mask a region and describe the product. Preview the placement before running the full-quality version.

🧹 Object Removal and Cleanup
Mask unwanted elements (wires, signage, background clutter) and prompt the model to fill with the surrounding environment. The inpainting LoRA is trained for clean region fills.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clear masks that cover the target region with some margin

  • Detailed prompts that describe the full scene including the new content

  • Single-object replacements on clean backgrounds

  • Footage with consistent lighting and a locked camera

⚠️ May produce softer results

  • Tiny masks on small details (the model needs room to work)

  • Prompts that describe only the new object without scene context

  • Fast-moving subjects that cross the mask boundary rapidly

  • Masks that cover most of the frame (generation is better than inpainting at that point)


FAQ

What is the difference between the 1-pass and multi-pass inpainting workflows?
This 1-pass version generates the inpainted content in a single sampling pass at native resolution. The multi-pass version adds two additional upscale refinement passes using the LTX spatial upscaler, producing sharper detail and cleaner edges. Use 1-pass for fast iteration, previews, and testing. Use multi-pass for final delivery.

What is video inpainting?
Video inpainting replaces or fills a masked region of a video with new generated content. You draw a mask over the area you want to change, describe what should appear there, and the model generates content that blends with the surrounding footage. This workflow uses LTX-Video 2.3 with a dedicated inpainting LoRA trained for masked region generation.

Why does the prompt matter so much for inpainting?
The model uses the prompt to understand what belongs in the masked region and how it relates to the rest of the scene. "Red umbrella" gives no context about lighting, perspective, or environment. "A woman walking through a sunny park holding a red umbrella, dappled light through trees" gives the model everything it needs to generate a convincing fill.

Can I remove an object without replacing it?
Yes. Mask the object and prompt the model to fill with the background. "A clean brick wall with afternoon sunlight" or "empty wooden table surface" tells the model to generate environment where the object was.

Does this workflow support audio?
Audio conditioning is included but bypassed by default. Enable it inside the workflow to generate scene-matched audio alongside the inpainted video. For most inpainting tasks, audio is not needed.

Is LTX-Video 2.3 free to use commercially?
Yes. LTX-Video 2.3 is released under the Apache 2.0 license, which allows commercial use, modification, and redistribution. You can use the outputs in client work, broadcast, published content, and commercial products.

How to run LTX 2.3 video inpainting online?
You can run LTX 2.3 video inpainting online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload a video, draw a mask, describe the change, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

An editor runs a quick inpaint test and likes the direction. A teammate opens that exact run from shared history and refines it with the multi-pass version. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Upload a video, draw a mask, describe the change, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N