AI Influencer Video Maker · Image to Video
Upload a person photo and a product photo, and this two-step workflow composites them into a styled influencer image with Qwen Image Edit 2511, then animates it into a 10-second talking video with synchronized audio using LTX-2 19B.
ai influencer
ltx2
opensource
product video
qwen image edit
1
57
Nodes & Models
LoadImage
CheckpointLoaderSimple
ltx-2-19b-distilled.safetensors
VAELoader
qwen_image_vae.safetensors
LatentUpscaleModelLoader
ltx-2-spatial-upscaler-x2-1.0.safetensors
CLIPLoader
qwen_2.5_vl_7b_fp8_scaled.safetensors
LTXVGemmaCLIPModelLoader
gemma-3-12b-it-qat-q4_0-unquantized/model-00001-of-00005.safetensors
ltx-2-19b-distilled.safetensors
UNETLoader
qwen_image_edit_2511_bf16.safetensors
LTXVAudioVAELoader
ltx-2-19b-distilled.safetensors
EmptyImage
PrimitiveFloat
Note
KSamplerSelect
ManualSigmas
RandomNoise
PrimitiveInt
OrchestratorNodeGroupBypasser
CLIPTextEncode
LoraLoaderModelOnly
Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors
ImageScaleBy
ImageScaleToTotalPixels
ModelSamplingAuraFlow
GetImageSize
TextEncodeQwenImageEditPlus
LTXVEmptyLatentAudio
EmptyLTXVLatentVideo
VAEEncode
CFGNorm
CreateVideo
FluxKontextMultiReferenceLatentMethod
LTXVAudioVAEDecode
CFGGuider
KSampler
SamplerCustomAdvanced
VAEDecode
ImpactExecutionOrderController
SaveImage
LTXVSeparateAVLatent
SaveVideo
LTXVConcatAVLatent
LTXVImgToVideoInplace
LTXVLatentUpsampler
LTXVConditioning
LTXVPreprocess
FloyoStickyNote
CM_FloatToInt
LTXVSpatioTemporalTiledVAEDecode
ResizeImagesByLongerEdge
CM_FloatToInt
ABOUT THE WORKFLOW
Create an Influencer Product Video
Upload two images: your model (the person) and your product. Step 1 composites them into a single lifestyle image of the person holding the product in a styled scene. Step 2 takes that composited image and animates it into a 10-second vertical video with spoken dialogue, natural motion, and synchronized audio. Both steps use open-source models. No partner node credits required.
Model
Qwen Image Edit 2511 (bf16) by Alibaba. Used in Step 1 to composite the person and product into one styled scene using multi-reference image editing, with the Lightning LoRA for 6-step generation.
LTX-2 19B Distilled by Lightricks. Used in Step 2 to animate the composited image into a talking video with native audio-visual generation, paired with a Gemma 3 12B text encoder and a two-pass spatial upscaling pipeline.
HOW IT WORKS
Step 1. Upload your model image
A photo of the person who will present the product. Front-facing, well-lit, visible face and upper body. Use a 9:16 (1080x1920) image for correctly sized output.
Works great with: portraits · influencer photos · AI-generated characters
Step 2. Upload your product image
A clear photo of the product on a clean background. Also 9:16 (1080x1920) for best results.
Works great with: sneakers · beauty products · gadgets · fashion accessories · packaged goods
Step 3. Run Step 1 to generate the composite image
Qwen Image Edit 2511 composites the person holding the product in the scene you describe. Preview the result before continuing.
Step 4. Load the generated image into Step 2
Upload or connect the output from Step 1 into the "Model Image With Product" input. Enable Step 2 in the workflow.
Step 5. Edit the video prompt and run Step 2
Describe the presenter's motion, dialogue, and camera framing. LTX-2 animates the image into a 10-second vertical video with lip-sync, breathing, blinking, and natural hand movement. Audio is generated alongside the video.
Ready for: TikTok · Instagram Reels · YouTube Shorts · ad platforms · e-commerce listings
First time? Run Step 1 first. Check the generated image. Then enable Step 2 and run again.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard influencer video — Run Step 1 with default settings (1.6MP, 6 steps). Run Step 2 with default settings (245 frames, 24fps, two-pass upscale). Edit the prompts to match your product.
Different setting or scene — Edit the Step 1 prompt. Replace "clean urban street with palm trees" with "minimalist kitchen counter," "studio with soft ring light," or any background that fits the product.
Different dialogue or presenter tone — Rewrite the video prompt in Step 2. Describe what the person says and how they move. Keep spoken lines short and conversational.
Multiple product variations — Keep the same model image and swap the product image across runs. Each run composites the person with a different product.
Different presenter — Swap the model image. Any front-facing portrait works, including AI-generated characters.
The person and product look disconnected — Edit the Step 1 prompt to describe the interaction more specifically. "He holds the sneaker at chest level, slightly angled toward the camera so the logo is visible" gives the model clear spatial cues.
Video motion looks stiff — Add more motion detail to the Step 2 prompt. "Subtle wrist adjustment while holding the product, slight head tilt, casual rotation of the item" gives the model specific physical actions to animate.
Prompt: Step 1 describes the scene composition. Step 2 describes motion and dialogue. "He casually rotates the sneaker to show the side profile and brings it closer to the camera while explaining" is specific. "Person shows product" is not.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
📱 UGC-Style Product Ads
Generate vertical talking-head product videos for TikTok and Reels by uploading a presenter photo and a product photo.
🛍️ E-commerce Product Listings
Create product demo videos for marketplace listings by swapping product images across multiple runs with the same presenter.
🧴 Beauty and Lifestyle Demos
Composite a model holding a skincare, fashion, or food product and generate a short review clip with natural motion and speech.
🔄 A/B Testing Ad Creatives
Produce multiple video variations with different presenters, products, or scripts to test which combination performs best before committing to production.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Clear, front-facing model photos with visible face and hands at 9:16
Product photos on clean backgrounds with strong lighting at 9:16
Specific motion descriptions in the Step 2 prompt (rotates, tilts, brings closer)
Simple scenes with one person and one product
⚠️ May produce softer results
Model photos with sunglasses, heavy occlusion, or extreme angles
Product photos with busy backgrounds or multiple items
Input images not in 9:16 aspect ratio (causes sizing issues)
Long monologue descriptions with no physical product interaction
FAQ
What models does this workflow use?
Step 1 uses Qwen Image Edit 2511 by Alibaba with the Lightning LoRA for fast 6-step image compositing. Step 2 uses LTX-2 19B Distilled by Lightricks with a Gemma 3 12B text encoder for video generation with native audio. Both are open-source models, so no external API credits are needed.
How is this different from the Nano Banana and Kling version of this workflow?
This version uses open-source models for both steps (Qwen Image Edit 2511 and LTX-2 19B) instead of partner-node APIs (Nano Banana Pro and Kling 2.6 Pro). You pay only for generation time on the GPU rather than per-run API credits. The trade-off is that generation takes longer due to the two-pass video pipeline with spatial upscaling.
Does the video include spoken dialogue with lip-sync?
Yes. Describe what the person says in the Step 2 prompt and LTX-2 generates voice, mouth movement, and ambient audio together as part of its native audio-visual generation.
Why does the workflow run in two steps instead of one?
Step 1 generates a static image where the person is holding the product in the correct pose and scene. You preview and approve this image before Step 2 animates it into video. This gives you control over the composition before committing to the longer video generation.
What resolution and duration does the video output have?
The video generates at 960x544 in the first pass, then upscales with a dedicated LTX spatial upscaler. The final output is approximately 10 seconds at 24fps in portrait format with synchronized audio.
Do the input images need to be a specific size?
Yes. Both the model image and the product image should be 9:16 aspect ratio (1080x1920) for correctly sized outputs. Images in other ratios may produce cropping or distortion issues.
How to run AI influencer video generation online?
You can run AI influencer video generation online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your inputs, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it?
Upload a model photo and a product photo, write the scene and dialogue prompts, and run both steps.
Questions? Watch the free course or check the FAQ above.
Read more
_1772389694491_1782979813618.webp?width=1400&height=620&quality=80&resize=cover)
_1772389694491_1782979813618.webp?width=104&height=104&quality=80&resize=cover)
_1772104822143_1782978375956.webp?width=400&height=300&quality=80&resize=cover)





