Wan 3.0 · Text to Video With Audio
Generate up to 30 seconds of 1080p video with sound using Wan 3.0, Alibaba's latest video model. Write a prompt, set the duration, and hit run. Audio included.
alibaba
text to video
tongyi lab
video with audio
wan 3.0
0
17
Nodes & Models
AlibabaWan30TextToVideo_floyo
VideoToFrames
CreateVideo
SaveVideo
ABOUT THE WORKFLOW
Build a 30-Second Clip From Text Write a prompt describing the scene, the action, and the camera. Wan 3.0 builds the picture and the audio together in one pass and returns a clip with sound. Up to 30 seconds in a single run at up to 1080p. No input image needed.
Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.
Model
Wan 3.0 by Tongyi Lab at Alibaba. Public beta 6 August 2026. The current flagship of the Wan video family. Unifies text-to-video, image-to-video, reference-to-video, and editing into a single model. Generates native 30-second clips at up to 1080p with audio in one pass. Includes intelligent duration control that matches clip length to the prompt. Closed model with no open weights.
HOW IT WORKS
Step 1. Write your prompt Describe the scene as one evolving motion. Name the subject, the camera path, the lighting, and the sound across the full take. Works great with: continuous shots · multi-character scenes · product spots · cinematic takes
Step 2. Set resolution and duration Pick from 480P to 1080P and a duration up to 30 seconds. Both settings drive what a run costs.
Step 3. Hit run and download The model builds the picture and the audio together and saves the clip under video/ComfyUI. Ready for: Premiere · DaVinci Resolve · CapCut · After Effects
First time? Leave every setting as-is. The defaults (1080P · 5 seconds · 16:9 · audio on · watermark off · random seed) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard clip (most people) — 1080P · 5 seconds · 16:9 · audio on · watermark off · random seed. The right starting point for almost everyone.
Need a longer clip — Raise duration up to 30 seconds. Write the prompt as one continuous take that fills the time, not a short idea padded out. Duration and resolution together drive cost.
Need a different shape — Switch the aspect ratio. 16:9, 9:16, and 1:1 are available.
Sound is not matching the action — Describe the sound separately in the prompt after the visual action. Name the room tone, effects, and score.
Want silent output — Turn audio off. The model skips the sound pass.
Repeat a take you liked — Set a fixed seed. The same prompt, shape, and seed give you the same clip back.
The clip is generic — Write a longer prompt. Name the camera speed, the lighting direction, the texture of surfaces, and the atmosphere. Short prompts return short clips that look like stock footage.
Prompt: Write it as one continuous take. "A woman walks through a night market, camera follows from behind, she pauses at a food stall, the vendor hands her a bowl, steam rises, she turns and smiles at the camera, neon signs reflect in puddles, ambient chatter and sizzling wok sounds." Wan 3.0 handles long prompts with evolving motion, multiple characters, and camera transitions in a single pass. Fill the duration with action, not filler.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 One-Take Scenes Generate a 30-second continuous shot with camera moves, character interactions, and ambient sound in a single pass.
📢 Ads and Product Spots Build a finished spot from a text brief at 1080p without a shoot, a storyboard, or a separate sound session.
📱 Social Content Turn a one-line idea into a ready-to-post clip with sound for Reels, TikTok, or Shorts. Up to 30 seconds from one prompt.
🎨 Previz and Concept Test how a scene plays at full length before committing to a production schedule or a camera day.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Detailed prompts written as one evolving take
Clips that describe action across the full duration
Multi-character scenes with named interactions
Prompts that describe both the picture and the sound
⚠️ May produce softer results
Short keyword prompts padded to 30 seconds
Expecting 4K output (the model caps at 1080P)
Fast collisions and extreme physics
Complex dialogue with precise phoneme-level mouth matching
FAQ
What is Wan 3.0? Wan 3.0 is the latest video model from Tongyi Lab at Alibaba, in public beta since 6 August 2026. It unifies text-to-video, image-to-video, reference-to-video, and editing into one model. It generates native 30-second clips at up to 1080p with audio in a single pass. Its signature features are 30-second single-take generation, intelligent duration control that matches clip length to the prompt, and Omni-Reference input that accepts documents and webpages alongside text, images, audio, and video. This workflow uses the text-to-video mode.
How long can a Wan 3.0 clip be? Up to 30 seconds in a single pass. That is double the 15-second ceiling on Wan 2.7 and matches Seedance 2.5 at the top of the current duration range. The model includes intelligent duration control that recommends a clip length based on your prompt.
Does Wan 3.0 generate audio with the video? Yes, by default. Sound is generated in the same pass as the picture, so effects, room tone, and ambient audio land on the same timeline. Describe the sound in your prompt and it goes onto the track. Turn audio off if you want a silent clip.
What is the difference between Wan 3.0 and Wan 2.7? Wan 2.7, released April 2026, goes up to 15 seconds and adds a thinking mode for compositional planning. Wan 3.0, beta August 2026, doubles the duration to 30 seconds, unifies four separate Wan 2.7 endpoints into one model, and adds Omni-Reference input that accepts documents and webpages. Neither has open weights.
Is Wan 3.0 open source? No. Wan 3.0 is API-only with no published weights. Open weights in the Wan family stop at Wan 2.2, which is released under Apache 2.0.
How does Wan 3.0 compare to Seedance 2.5 and Kling 3.0 Pro? All three generate video with native audio. Wan 3.0 and Seedance 2.5 share the 30-second ceiling. Kling 3.0 Pro leads on resolution (native 4K) and frame rate (60 fps) but caps at 15 seconds. Wan 3.0 adds Omni-Reference document input. Seedance 2.5 supports up to 50 multimodal references. Pick Wan 3.0 for unified generation and editing in one model, Seedance 2.5 for the highest reference count, and Kling 3.0 Pro for 4K output.
How to run Wan 3.0 online? You can run Wan 3.0 online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, write a prompt, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it? Write a prompt and run it. The clip comes back with sound, up to 30 seconds.
Questions? Watch the free course or check the FAQ above.
Read more



