Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

Wan 2.5: Image to Video with Audio

27.1k

Generates in about 2 mins 7 secs

Nodes & Models

AlibabaWan25ImageToVideo_floyo
VideoToFrames
LoadImage
VHS_VideoCombine

ABOUT THE WORKFLOW

Animate a Photo With Sound Upload an image and describe the motion you want. Wan 2.5 writes the picture and the matching sound in one pass and returns a short clip. Point it at your own audio file to drive the sound, or leave that field empty and let the model write its own.

Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.

Model

  • Wan 2.5 by the Wan team at Alibaba. An audio-video model that generates picture, speech with lip sync, and sound effects together rather than in separate passes.


HOW IT WORKS

Step 1. Upload your image The picture you want to animate. It sets the opening frame, the framing, and the look. Works great with: portraits · product shots · characters · illustrations

Step 2. Describe the motion Say what moves and how, not what is already in the frame. "The cat leans in and sniffs the flowers, petals sway, slow push in."

Step 3. Add an audio track (optional) Paste a link to a sound file and the model matches the clip to that track. Leave it empty and the model writes audio from your prompt instead.

Step 4. Set resolution and length Pick 480P, 720P, or 1080P, and 5 or 10 seconds. Both settings drive what a run costs.

Step 5. Hit run and download Wan 2.5 returns the clip and the workflow saves it as an MP4 at 24 frames per second. Ready for: Premiere · DaVinci Resolve · CapCut · After Effects

First time? Leave every setting as-is. The defaults (720P · 5 seconds · sound on · prompt extend on · random seed) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard clip (most people) — 720P · 5 seconds · sound on · prompt extend on · random seed. The right starting point for almost everyone.

  • Testing a prompt first — Drop to 480P at 5 seconds. Cheapest way to see whether the motion reads before you pay for a full run.

  • Final delivery or client work — 1080P at 5 or 10 seconds. Resolution and length both drive cost, so save this for the take you plan to keep.

  • Want speech or a specific track — Paste a link to your audio file into the audio field. The model matches the clip to that track instead of writing its own.

  • Repeat a take you liked — Set a fixed seed number instead of leaving it on random. The same image, prompt, and seed give you the same clip back.

  • The motion is wrong — Rewrite the prompt before you touch a setting. Name the movement and the camera move, like "slow push in" or "she turns her head to the left."

  • The same artifact keeps appearing — The negative prompt already covers common failures. Add the one you keep seeing, such as "warped hands" or "flicker."

  • Your prompt is one short line — Leave prompt extend on. It rewrites a thin prompt into a fuller one before the clip is generated.

Prompt: Describe motion, not contents. "The woman turns toward the window and smiles, curtains drift, slow push in" gives you more than "a woman in a room." If you want specific sound, name it in the prompt: "footsteps on gravel, wind in the trees."


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Social Video Turn one product shot or portrait into a short clip with matching sound for Reels, TikTok, or Shorts.

🗣️ Talking Characters Feed a character portrait and a voice track, and get a clip with the mouth matched to the words.

📢 Ads and Promos Animate a still product image into a 5 or 10 second spot without booking a shoot.

🎨 Concept and Previz Test how a still frame moves before committing to a full animation pass or a camera day.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clear, well-lit source images with one main subject

  • Prompts that name the movement and the camera move

  • Steady camera work like a slow push in or a pan

  • Short clips of 5 seconds

⚠️ May produce softer results

  • Fast, tangled motion with several subjects at once

  • Prompts that describe the image instead of the motion

  • Low-resolution or heavily compressed source images

  • Dialogue that runs longer than the clip


FAQ

What is Wan 2.5? Wan 2.5 is a video model from the Wan team at Alibaba, released in September 2025 as Wan 2.5 Preview. It takes a text prompt, a still image, and an optional audio track, then writes picture and sound in one pass, including speech with lip sync, sound effects, and background sound. It runs at 480p, 720p, and 1080p, 24 frames per second, in 5 or 10 second clips.

Does Wan 2.5 generate audio with the video? Yes. The sound is written in the same pass as the picture, so dialogue and effects line up with the action instead of being added afterwards. In this workflow the returned clip is rebuilt into an MP4 before saving, and that saved file carries picture only. Connect the audio output on the frame step into the audio input on the video save node to keep the generated sound in the file.

Is Wan 2.5 open source? No. Wan 2.5 shipped as an API-only model and the weights were never published to Hugging Face, GitHub, or ModelScope. Wan 2.1 and Wan 2.2 are the open-weight releases in the family. If you need weights you can download and run on your own hardware, use Wan 2.2.

What is the difference between Wan 2.5 and Wan 2.2? Wan 2.2 is open weight, runs locally, and generates picture only. Wan 2.5 runs through an API and adds synchronized audio, resolution up to 1080p, and clips up to 10 seconds. Pick Wan 2.2 for local control with no per-run cost, and Wan 2.5 when you want sound generated with the picture.

Can I use my own audio track with Wan 2.5? Yes. The audio field takes a link to a sound file, and the model matches the clip to that track rather than writing its own. Leave the field empty and Wan 2.5 generates audio from your prompt.

How long can a Wan 2.5 clip be? Five or ten seconds per run, at 24 frames per second. Longer pieces are built by running several clips and joining them in an editor. Both length and resolution affect what a run costs.

How to run Wan 2.5 online? You can run Wan 2.5 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your image, describe the motion, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Upload an image, describe the motion, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N
l
leonardopedroso
6 months ago
criar imagem do perigo de uma criança sozinha na rua

Reply

n
nightlucky
6 months ago
Crie um video meio anime de um garoto fantasma que possui o corpo de uma jovem em uma festa

Reply