Wan 2.7 · Image to Video With Audio
Generate 1080p video with sound using Wan 2.7, Alibaba's flagship video model with a thinking mode. Upload a start image, describe the action, hit run.
animation
film production
image to video
Video
video generation
wan
4
1.5k
Nodes & Models
AlibabaWan27ImageToVideo_floyo
VideoToFrames
LoadImage
CreateVideo
SaveVideo
ABOUT THE WORKFLOW
Animate a Still With Sound Upload a picture and describe the action you want. Wan 2.7 plans the composition before it renders, then builds the picture and the audio in one pass. Add an optional last frame to steer where the shot ends. The clip comes back as a video file with sound.
Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.
Model
Wan 2.7 by Tongyi Lab at Alibaba. Released April 2026. The current flagship of the Wan video family. A 27 billion parameter mixture-of-experts diffusion transformer that generates 720p or 1080p video up to 15 seconds with native audio in one pass. Features a thinking mode that plans the composition before the model renders.
HOW IT WORKS
Step 1. Upload your start image The frame the clip opens on. It sets the subject, the scene, and the lighting. Works great with: portraits · product shots · street scenes · concept art
Step 2. Describe the action Say what moves and how. Be specific about motion, camera, and mood. "The three dancers step side to side in sync, hair moves with each turn, camera holds steady."
Step 3. Upload a last frame (optional) The frame the clip should land on. Helps the model plan the motion path from start to finish.
Step 4. Set resolution and length Pick 720P or 1080P and a duration up to 15 seconds. Both settings drive what a run costs.
Step 5. Hit run and download The clip comes back as a video file with sound, saved under video/Wan2.7. Ready for: Premiere · DaVinci Resolve · CapCut · After Effects
First time? Leave every setting as-is. The defaults (1080P · 6 seconds · prompt extend on · watermark off · random seed) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard clip (most people) — 1080P · 6 seconds · prompt extend on · watermark off · random seed. The right starting point for almost everyone.
Quick test before a full run — Drop to 720P at 6 seconds. Cheaper way to check whether the motion and the sound read the way you wanted.
Need a longer clip — Raise duration up to 15 seconds. Duration and resolution together drive cost, so raise one at a time.
Want the shot to land on a specific frame — Upload a second image as the last frame. The model plans the motion between the two.
Repeat a take you liked — Set a fixed seed number instead of leaving it on random. The same image, prompt, and seed give you the same clip back.
Your prompt is already detailed — Turn prompt extend off. Left on, the model rewrites your prompt for richer detail before generating, which helps loose descriptions and works against carefully worded ones.
Motion is not following your prompt — Rewrite the prompt before you touch a setting. Name each movement, the camera behaviour, and what stays fixed.
The negative prompt — The default lists common quality issues. Add anything you keep seeing, like "warped hands" or "flickering."
Prompt: Write the action, not the contents. "She leans forward, picks up the cup, steam rises, camera pushes in from the side" gives you more than "a woman with a cup." Name the camera move separately from the subject, and say what stays fixed. If you want specific sound, name it: "espresso machine hissing, low cafe murmur."
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 Social Video Turn one still into a short clip with matching sound for Reels, TikTok, or Shorts. Up to 15 seconds in a single run.
📢 Ads and Product Films Animate a product shot into a finished spot at 1080p without a shoot or a separate sound pass.
🎞️ Start-to-End Shots Give the model a first frame and a last frame and get controlled motion between the two with matching audio.
🎨 Previz and Storyboarding Test how a frame moves and sounds before committing to a full production pass or a camera day.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Clear, well-lit start images with one readable subject
Prompts that name the movement, the camera, and the sound
First and last frame pairs that share similar framing
Clips of 6 to 10 seconds
⚠️ May produce softer results
Prompts that describe the image instead of what happens next
Fast collisions or complex physics
Small text that needs to stay readable through motion
Duration pushed to 15 seconds on a first test
FAQ
What is Wan 2.7? Wan 2.7 is the flagship video model from Tongyi Lab at Alibaba, released April 2026. It is a 27 billion parameter mixture-of-experts diffusion transformer that generates 720p or 1080p video up to 15 seconds with native audio in one pass. Its standout feature is a thinking mode that plans the composition before the model renders, which improves character consistency, motion adherence, and narrative coherence.
Does Wan 2.7 generate audio with the video? Yes. Sound is generated in the same pass as the picture, including ambient sound, effects, and dialogue with lip sync. Describe the sound in your prompt and it goes onto the track.
Is Wan 2.7 open source? No. Wan 2.7 is API-only with no published weights. Open weights in the Wan family stop at Wan 2.2, which is released under Apache 2.0. Pick Wan 2.2 for open weights and local control, and Wan 2.7 for the latest quality, thinking mode, and native audio through the API.
What is the difference between Wan 2.7 and Wan 2.5? Wan 2.5 was the first in the family to add native audio and shipped as a preview in September 2025, with clips up to 10 seconds. Wan 2.7 extends that to 15 seconds, adds a thinking mode that plans composition before rendering, and improves character consistency and first-to-last-frame control. Neither has published weights.
What is the difference between Wan 2.7 and Wan 2.2? Wan 2.2 is open weight under Apache 2.0, runs on your own hardware, and generates picture only. Wan 2.7 runs through an API, adds native audio, a thinking mode, and first-to-last-frame interpolation. Pick 2.2 for local control and no per-run cost, and 2.7 for the latest capabilities through the endpoint.
Can Wan 2.7 interpolate between a start and end frame? Yes. Upload a second image as the last frame and the model plans the motion between the two. Both images should frame the subject at a similar size for the cleanest result.
How to run Wan 2.7 online? You can run Wan 2.7 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your image, describe the action, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it? Upload a start image, describe the action, and run it. The settings are already set.
Questions? Watch the free course or check the FAQ above.
Read more




