Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

Wan 2.1 FusionX · Image to Video

Animate a still into a cinematic clip with Wan 2.1 FusionX, a community fine-tune of Alibaba's 14B video model. Upload your image, describe the shot, hit run.

5.2k

Generates in about 2 mins 38 secs

Nodes & Models

LoadImage
WanVideoBlockSwap
WanVideoImageToVideoEncode
WanVideoTorchCompileSettings
WanVideoClipVisionEncode
CLIPVisionLoader
ImageResizeKJv2
LoadWanVideoT5TextEncoder
WanVideoDecode
WanVideoVAELoader
WanVideoModelLoader
WanVideoSampler
WanVideoTextEncode
VHS_VideoCombine

ABOUT THE WORKFLOW

Turn a Still Into a Cinematic Clip
Upload a reference image and describe the shot. FusionX renders 81 frames with a cinematic bias in 10 passes and saves them as an MP4 at 16 frames per second. Add a second reference image for blended compositions, or leave it off for a single-image run.

Model

  • Wan 2.1 FusionX by Vrgamedevgirl84. A community fine-tune of Alibaba's Wan 2.1 14B image-to-video model, merged with CausVid, AccVideo, and MoviiGen to push motion quality, scene consistency, and visual detail. Tuned for cinematic camera moves, dramatic lighting, and fast iteration at 10 steps.


HOW IT WORKS

Step 1. Upload your reference image
The image the clip starts from. A clear subject with cinematic framing works best.
Works great with: character art · landscapes · product shots · concept frames

Step 2. Describe the shot
Say what happens and how the camera moves. Short and specific beats long and vague. "Cinematic shot of a dog running through a forest."

Step 3. Enable a second reference (optional)
A second image loader sits muted by default. Enable it and upload a second reference for blended or composite shots.

Step 4. Hit run and download
The model renders 81 frames, decodes them, and saves a 5-second MP4 at 16 fps under fusionX_Image2Video/video.
Ready for: Premiere · DaVinci Resolve · CapCut · After Effects

First time? Leave every setting as-is. The defaults (1280 x 720 · 81 frames · 10 steps · guidance 2 · fixed seed) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard clip (most people) — 1280 x 720 · 81 frames · 10 steps · guidance 2 · fixed seed. The right starting point for almost everyone.

  • Motion is jittery — Lower guidance from 2 toward 1. The model follows the prompt less tightly but the motion smooths out.

  • The clip ignores the prompt — Raise guidance one point at a time. Higher values hold the model closer to your words.

  • Faster test runs — Lower steps to 6 or 8. FusionX is merged from speed-tuned models, so it holds up well at lower step counts.

  • Want a different take — Change the seed number. The default runs on fixed, so the same image, prompt, and seed return the same clip.

  • Want a blended or composite shot — Enable the second image loader (right-click, Unmute or Ctrl+M) and upload a second reference. The model blends both into the clip.

  • Resolution above 720p — The model was fine-tuned at 768 square and tested at 1280 x 720. Going larger costs memory and time without a guaranteed quality gain. The creator recommends shifting guidance from 1 to 2 when moving from 1024 x 576 to 1280 x 720.

Prompt: Keep it short and cinematic. "Cinematic close-up of a samurai drawing a katana in heavy rain, camera slowly orbits left, lightning flash" gives you more than "samurai in rain." Name the camera move, the lighting, and the mood. The negative prompt is in Chinese and covers common quality issues; leave it as-is.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Cinematic Shots
Turn concept art or stills into clips with dramatic camera moves, volumetric lighting, and moody atmospherics.

🎨 Concept and Previz
Test how a frame moves before committing to a full production pass. Ten steps finish fast enough to iterate.

🖼️ Dual-Reference Compositing
Blend two reference images into one clip for character-in-environment shots or style transfers.

📱 Social and Portfolio Content
Turn one still into a short, polished clip for Reels, TikTok, or a creative portfolio.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clear, well-lit reference images with cinematic framing

  • Short prompts that name the camera move and the mood

  • Step counts of 6 to 10

  • 1280 x 720 or 768 x 768 resolution

⚠️ May produce softer results

  • Readable text in the scene

  • Resolutions pushed past the 720p training range

  • Long descriptive prompts that overload the model

  • Complex multi-character choreography


FAQ

What is Wan 2.1 FusionX?
Wan 2.1 FusionX is a community fine-tune of Alibaba's Wan 2.1 14B image-to-video model, created by Vrgamedevgirl84. It merges several research-grade models, including CausVid, AccVideo, and MoviiGen, to improve motion quality, scene consistency, and visual detail. The result renders up to 50% faster than the base Wan 2.1 model and holds cinematic quality at as few as 6 to 8 steps.

Is Wan 2.1 FusionX free for commercial use?
Not straightforwardly. The base Wan 2.1 model is Apache 2.0, but FusionX is a merge of multiple components, and some (like CausVid) are released under CC BY-NC-SA 4.0, which restricts commercial use. The creator states that commercial use is not permitted for models or components under non-commercial licenses. Do your own legal due diligence before using output commercially.

How does FusionX compare to the base Wan 2.1 14B?
FusionX is a drop-in replacement that generates at the same resolution and frame count but with stronger motion, richer detail, and faster convergence. The creator reports up to 50% faster generation with SageAttn enabled. The trade-off is the mixed licensing from the merged components.

Why is the negative prompt in Chinese?
Wan 2.1 was trained on bilingual data, and the Chinese negative prompt targets quality issues more precisely in the model's training language. It covers oversaturation, static frames, blurred detail, deformed limbs, and cluttered backgrounds. Leave it as-is.

Can I use two reference images at once?
Yes. A second image loader sits muted by default. Enable it and upload a second reference, and the model blends both into the clip. This works for character-in-environment composites and style transfers.

What resolution and length does this workflow produce?
It ships at 1280 x 720, 81 frames at 16 fps, which gives about 5 seconds. The model was fine-tuned at 768 square. Going larger costs memory without a guaranteed quality gain.

How to run Wan 2.1 FusionX online?
You can run Wan 2.1 FusionX online through Floyo. No installation, no setup, no model downloads. Open the workflow in your browser, upload a reference image, describe the shot, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Upload a reference image, describe the cinematic shot, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N