Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing
Run now Wan3.0 on Floyo hero

COMMUNITY PAGE

Run now Wan3.0 on Floyo

Home / Model / Wan 3.0 on Floyo

AI VIDEO GENERATION

Run Wan 3.0 on Floyo

Alibaba's flagship video model with native 30-second single-shot generation. Feed it text, images, video, audio, documents (PDF/Excel/Word/PowerPoint), or web URLs. Smart duration control. Pixel-level reference consistency. Two tiers: Standard and Prime. Up to 1080p with native synchronized audio.

Run Alibaba's Wan 3.0 through ComfyUI in your browser. No API key, no installs, no local GPU.

Duration

Up to 30 seconds (single shot)

New Input

Documents + web URLs

Resolution

480p / 720p / 1080p

Tiers

Standard (~14 min) + Prime (~2 min)

No installation. Runs in browser. Updated August 2026.

What to get?

Wan 3.0 is Alibaba's flagship video model, launched in public beta August 6, 2026. Native 30-second single continuous shot generation (double Wan 2.7's 15-second ceiling). Six input modalities: text, images, video, audio, documents (PDF/Excel/Word/PowerPoint/TXT), and web page URLs. Smart duration control matches clip length to content. Pixel-level consistency for characters, objects, scenes, and styles across the full clip. Higher realism with natural micro-expressions, faithful color reproduction, and multilingual voice. Two tiers on Floyo: Standard (wan3.0-video, ~14 min at 720p/15s, from $0.05/s) and Prime (wan3.0-video-prime, ~2 min at 720p/15s, from $0.068/s). Available as ComfyUI API nodes on Floyo with 6 workflows.

WAN 3.0 WORKFLOWS ON FLOYO

Wan 3.0 Text to Video with Audio

Wan 3.0 Image to Video with Audio

Wan 3.0 Reference to Video

Wan 3.0 Prime Text to Video with Audio

Wan 3.0 Prime Image to Video with Audio

Wan 3.0 Prime Reference to Video

alibaba

image to video

video with audio

wan 3.0 prime

Animate a still into a high-fidelity clip with sound using Wan 3.0 Prime, the quality tier of Alibaba's latest video model. Upload an image, describe the shot.

Wan 3.0 Prime · Image to Video With Audio

Animate a still into a high-fidelity clip with sound using Wan 3.0 Prime, the quality tier of Alibaba's latest video model. Upload an image, describe the shot.

alibaba

image to video

video with audio

wan 3.0

Animate a still into a 30-second clip with sound using Wan 3.0, Alibaba's latest video model. Upload an image, describe the motion, and hit run. 1080p output.

Wan 3.0 · Image to Video With Audio

Animate a still into a 30-second clip with sound using Wan 3.0, Alibaba's latest video model. Upload an image, describe the motion, and hit run. 1080p output.

multi-reference

reference to video

video to video

wan 3.0 prime

Place a character into any video using Wan 3.0 Prime by Alibaba. Upload a face, a reference clip, and describe the scene. Prime fidelity, audio included.

Wan 3.0 Prime · Reference to Video

Place a character into any video using Wan 3.0 Prime by Alibaba. Upload a face, a reference clip, and describe the scene. Prime fidelity, audio included.

alibaba

video style transfer

video to video

wan 3.0

Re-render any video in a new visual style using Wan 3.0 by Alibaba. Upload a clip, describe the target look, and hit run. Motion preserved, audio included.

Wan 3.0 · Reference to Video

Re-render any video in a new visual style using Wan 3.0 by Alibaba. Upload a clip, describe the target look, and hit run. Motion preserved, audio included.

alibaba

text to video

tongyi lab

video with audio

wan 3.0

Generate up to 30 seconds of 1080p video with sound using Wan 3.0, Alibaba's latest video model. Write a prompt, set the duration, and hit run. Audio included.

Wan 3.0 · Text to Video With Audio

Generate up to 30 seconds of 1080p video with sound using Wan 3.0, Alibaba's latest video model. Write a prompt, set the duration, and hit run. Audio included.

alibaba

text to video

video with audio

wan 3.0 prime

Generate a high-fidelity clip with sound from a written prompt using Wan 3.0 Prime, the quality tier of Alibaba's latest video model. No image needed, just describe the shot.

Wan 3.0 Prime · Text to Video With Audio

Generate a high-fidelity clip with sound from a written prompt using Wan 3.0 Prime, the quality tier of Alibaba's latest video model. No image needed, just describe the shot.

What is Wan 3.0?

Wan 3.0 is Alibaba's flagship video model, launched in public beta on August 6, 2026 through Alibaba Cloud Model Studio and QwenCloud. It generates up to 30 seconds of continuous video in a single shot with synchronized audio. This is double the 15-second ceiling of Wan 2.7. The 30-second duration enables real camera language: a push-in, a pan across a scene, or a tracking move can play out uninterrupted rather than resetting every few seconds.

The headline new capability is expanded input modalities. Beyond text, images, video, and audio, Wan 3.0 accepts documents (PDF, Excel, Word, PowerPoint, TXT) and web page URLs as generation references. Point the model at a product page, an article, or a research paper and it uses that content as reference material for the generated video. Upload a spreadsheet and it builds a dynamic chart presentation video. Upload a slide deck and it creates an animated introductory short.

Smart duration control lets the model suggest an appropriate clip length based on the described action and pacing. If your prompt describes a 5-second moment, the model does not pad it to fill 30 seconds. If it describes a complex sequence, the model takes the time it needs. You can also set exact duration manually. An extend function lets you continue any generated clip to build longer sequences from a coherent base.

Two tiers serve different production needs. Standard (wan3.0-video) generates a 720p/15-second clip in about 14 minutes and costs $0.10 per second. Prime (wan3.0-video-prime) generates the same clip in about 2 minutes at $0.14 per second. Both tiers support 480p, 720p, and 1080p. Use Standard for overnight batch work. Use Prime when you need results in your meeting.

On Floyo, Wan 3.0 runs through ComfyUI API nodes on H100 NVL GPUs. Six workflows cover text-to-video, image-to-video, and reference-to-video for both Standard and Prime tiers. No Alibaba Cloud account, no QwenCloud setup, no API key management.

What are Wan 3.0's technical specifications?

Wan 3.0 generates up to 30 seconds of continuous video with native synchronized audio in a single shot. Six input modalities: text, images, video, audio, documents, and web URLs. Three resolution tiers: 480p, 720p, 1080p. Two model tiers: Standard (~14 min for 720p/15s) and Prime (~2 min for 720p/15s). Smart duration control. Extend function for longer sequences. Pixel-level reference consistency. API-only (no open weights).

Spec Details
DeveloperAlibaba (Tongyi Lab)
Duration2-30 seconds per generation (single continuous shot)
Resolution480p, 720p, 1080p
AudioNative synchronized (dialogue, ambient, SFX, multilingual voice)
Smart DurationModel suggests appropriate clip length based on content and pacing
ExtendContinue any generated clip to build longer sequences
INPUT MODALITIES
TextImproved multi-part instruction following
Images / Video / AudioReference material for characters, style, motion, voice
Documents (New)PDF, Excel, Word, PowerPoint, TXT as generation references
Web URLs (New)Product pages, articles, papers, marketing sites as references
MODEL TIERS
wan3.0-video (Standard)~14 min for 720p/15s. Pricing: $0.05/s (480p), $0.10/s (720p), $0.20/s (1080p)
wan3.0-video-prime (Prime)~2 min for 720p/15s. Pricing: $0.068/s (480p), $0.14/s (720p), $0.28/s (1080p)
PLATFORM
Reference ConsistencyPixel-level across characters, objects, audio, scenes, and styles
Open SourceNo (API-only). Wan 2.2 is the last open-weight Wan model.
ComfyUI AccessAPI-based nodes on Floyo (6 workflows: 3 Standard + 3 Prime)
Release DateAugust 6, 2026 (public beta)

What can you create with Wan 3.0?

Wan 3.0 covers long-form single-shot video generation (up to 30 seconds), document-to-video conversion, web page-to-video conversion, product showcase ads, data visualization videos, animated presentations, character-consistent brand content, multilingual voiceover, and reference-based style and motion transfer. The 30-second duration and document input open production scenarios that shorter, prompt-only models cannot handle.

Capability What It Does Use Case
30-Second Single ShotGenerate up to 30 seconds of continuous video in one pass. Real camera language (push-in, pan, tracking) plays out uninterrupted. No clip stitching. Smart duration auto-adjusts to content.Broadcast spots, product walkthroughs, one-take narratives
Document-to-VideoUpload a PDF, spreadsheet, Word doc, or slide deck. The model reads the content and builds a video from the data, text, and structure inside. A spreadsheet becomes a dynamic chart video. A slide deck becomes an animated intro.Report presentations, data stories, sales decks, training content
Web URL-to-VideoPoint the model at a product page, article, or marketing site. It extracts content, design, and context, then generates a video that represents the page. A software website becomes a UI walkthrough promo. A product listing becomes a showcase ad.Product launch videos, web content repurposing, link-to-ad
Reference ConsistencyPixel-level consistency for characters, objects, audio, scenes, and styles across the full 30-second clip. A brand character looks the same at second 1 and second 28.Brand campaigns, series content, character-led stories
Native Audio + Multilingual VoiceSynchronized dialogue, ambient sound, SFX, and multilingual voice generated in the same pass. Natural emotional expression with micro-expressions synced to body movement.Talking head content, multilingual ads, immersive scenes
Pipeline IntegrationChain with image models in ComfyUI. Generate a character with Nano Banana or Ideogram V4, animate with Wan 3.0, add custom narration with ElevenLabs, upscale with Topaz. All in one workflow.Multi-model production pipelines

What are Wan 3.0's key features?

Wan 3.0's feature set is defined by two shifts: longer clips and wider inputs. The 30-second single shot doubles the industry standard. Document and URL inputs open modalities no other video model accepts. Everything else (better realism, stronger consistency, improved instruction following) compounds these two changes into a model built for professional content production at scale.

Native 30-Second Single Shot

Generate a full 30-second continuous clip in one pass. Previous Wan models capped at 15 seconds. Most competitors cap at 5-16 seconds. A 30-second clip covers a broadcast spot in a single generation with no cuts to stitch or match. Camera movements (push-in, pan, tracking, crane) play out with sustained coherence. Smart duration control auto-adjusts: a 5-second action generates at 5 seconds, not padded to 30. The extend function continues any clip for even longer sequences.

Document-to-Video

Upload a PDF, Excel spreadsheet, Word document, PowerPoint deck, or TXT file. The model reads the content and builds video from it. A data table becomes a dynamic chart presentation. A text introduction becomes an animated explainer. A slide deck becomes a polished introductory short. This is a modality no other video model supports. It collapses the "read the document, write a script, generate the video" pipeline into a single step.

Web URL-to-Video

Point the model at a product page, software website, research paper, or marketing site. It extracts the content, design patterns, and context, then generates a video that represents the page. A product listing becomes a showcase ad. A software landing page becomes a UI walkthrough promo. A research paper becomes an animated summary. This is the second new modality that Wan 3.0 introduces.

Pixel-Level Reference Consistency

Characters, objects, audio, scenes, and styles stay stable across the full generated sequence. This is described as pixel-level consistency. A product's color, texture, and proportions hold from the first frame to the last. A character's face, clothing, and body language stay recognizable throughout a 30-second continuous take. Previous Wan versions had consistency, but 3.0 pushes it to a fidelity level that holds up under close inspection.

Two Tiers: Standard and Prime

Standard (wan3.0-video) takes about 14 minutes for a 720p/15-second clip at $0.10 per second. Prime (wan3.0-video-prime) takes about 2 minutes for the same output at $0.14 per second. Both produce the same quality. The difference is speed. Use Standard for batch generation and overnight rendering. Use Prime when you need results during a client call or creative session. Both support 480p, 720p, and 1080p.

Higher Realism and Expressiveness

Visual details are closer to real footage. Portraits are distinctive, not generic. Color reproduction is faithful. Emotional expression includes micro-expressions synchronized to body movement. Multilingual voice response sounds natural across multiple languages and dialects. Text rendering in video is improved over previous versions. Motion and audio expressiveness are both stronger.

How does Wan 3.0 compare to other video models?

Wan 3.0 leads on clip duration (30 seconds native) and input modality breadth (documents + URLs + media). HappyHorse 1.0 leads on arena ranking (Elo 1,392) and cinematic depth-of-field. MiniMax H3 leads on reference input capacity (9 images + 3 videos + 3 audio). Gemini Omni Flash leads on stateful conversational editing. Wan 3.0's edge: the longest single-shot generation, document and URL inputs, and two speed tiers for different production budgets.

Model Max Duration Doc/URL Input Native Audio Max Resolution
Wan 3.0 30 seconds Yes (PDF/XLS/DOC/PPT + URLs) Yes 1080p
MiniMax H3 15 seconds No Yes Native 2K
HappyHorse 1.0 15 seconds No Yes (7-lang lip-sync) 1080p
Gemini Omni Flash ~10 seconds No Yes 720p (reported)
Wan 2.7 15 seconds No No Up to 4K

Source: Alibaba Wan 3.0 feature highlights document, Alibaba Cloud Model Studio pricing, QwenCloud wan3.0-video API reference, IT之家 and 第一财经 launch coverage (August 6, 2026), Pexo explainer, Siray Blog analysis, and Morphic model card.

How does Wan 3.0 work?

Wan 3.0 processes six input modalities (text, images, video, audio, documents, and web URLs) and generates video with synchronized audio in a single pass. The model understands content from each input type: text from documents, design from web pages, motion from reference video, voice from reference audio, and identity from reference images. It fuses these into a coherent audiovisual output.

For document inputs, the model reads the structure, data, and text from the uploaded file. A spreadsheet with quarterly revenue becomes a dynamic chart animation. A PowerPoint deck becomes an animated walkthrough. The model interprets the document's content semantically, not as a screenshot. For web URL inputs, the model crawls the page, extracts content, layout, and visual style, and generates video that represents the page's message and design.

Smart duration control works during the planning stage. Given a prompt describing a 5-second action, the model generates 5 seconds. A prompt describing a complex 30-second sequence gets the full duration. You can override this by setting duration manually. The extend function continues any generated clip from its last frame, maintaining visual and audio continuity.

On Floyo, Wan 3.0 runs through ComfyUI API nodes on H100 NVL GPUs. Six workflows cover both Standard and Prime tiers across text-to-video, image-to-video, and reference-to-video modes. Your inputs are sent to Alibaba's inference servers, and the generated video with audio returns as MP4 to your ComfyUI canvas. Chain with other models for extended pipelines.

Fair warning: Wan 3.0 is API-only. No open weights have been released. Alibaba's last open-weight Wan model is Wan 2.2 (Apache 2.0, July 2025). Alibaba has not published a technical report, parameter count, or architecture disclosure for Wan 3.0. The Standard tier takes about 14 minutes per 720p/15s clip, which is slow for interactive work (use Prime for speed). Content filtering is active. API pricing applies through your Floyo API Wallet. Prime costs roughly 40% more than Standard for the same output.

Frequently Asked Questions

Common questions about running Wan 3.0 on Floyo.

Is Wan 3.0 free to use on Floyo?

You can start with Floyo's free pricing plan. Floyo gives $0.25 in free API credits on signup. To continue using the service beyond the free tier, upgrade your Floyo pricing plan. Wan 3.0 runs as an API node, so generation costs come from your API Wallet. Standard starts at $0.05/s (480p). Prime starts at $0.068/s (480p).

How do I run Wan 3.0 without installing anything?

Open Floyo in your browser, search "Wan 3.0" in the template library, and pick a workflow (Standard or Prime, T2V, I2V, or R2V). Click Run, write your prompt, optionally upload reference media, and generate. Floyo handles the ComfyUI environment and API connection. No Alibaba Cloud account, no API key management.

Who made Wan 3.0?

Alibaba's Tongyi Lab. Wan 3.0 entered public beta on August 6, 2026 through Alibaba Cloud Model Studio and QwenCloud. It is the successor to Wan 2.7 (thinking mode, 4K, open source) and represents the first full architectural revision in the Wan series. Wan 3.0 is API-only with no open weights.

What is the difference between Wan 3.0 Standard and Prime?

Same output quality, different speed. Standard (wan3.0-video) takes about 14 minutes for a 720p/15-second clip at $0.10/s. Prime (wan3.0-video-prime) takes about 2 minutes for the same clip at $0.14/s. Both support 480p, 720p, and 1080p. Use Standard for batch work. Use Prime when you need results fast.

Can I upload a PDF or spreadsheet to generate video?

Yes. This is a new capability in Wan 3.0. Upload PDF, Excel, Word, PowerPoint, or TXT files as generation references. The model reads the content and builds video from the data and text inside. A spreadsheet becomes a dynamic chart presentation. A slide deck becomes an animated introductory video.

How does Wan 3.0 compare to Wan 2.7?

Wan 3.0 doubles the duration (30 seconds vs 15), adds document and URL inputs, adds native synchronized audio, and improves reference consistency and realism. Wan 2.7 offers thinking mode, up to 4K resolution, open-source weights (Apache 2.0), and LoRA training. Wan 3.0 is API-only. Wan 2.7 is open source. Both are available on Floyo.

Can I combine Wan 3.0 with other AI models in one workflow?

Yes. Floyo runs ComfyUI, which lets you chain multiple models. Generate a character with Nano Banana or Ideogram V4, animate with Wan 3.0, add narration with ElevenLabs or Fish Audio S2, upscale with Topaz. All in one pipeline.

Is Wan 3.0 open source?

No. Wan 3.0 is API-only with no published weights. Alibaba's last open-weight Wan model is Wan 2.2 (Apache 2.0, July 2025). Wan 2.7 also has open-source weights under Apache 2.0. If self-hosting and open weights are required, use Wan 2.7 on Floyo. For the latest capabilities (30-second clips, document input, URL input), use Wan 3.0 via the API.

Try Wan 3.0 on Floyo

Alibaba's flagship video model. 30-second single-shot generation, document and URL inputs, native audio, pixel-level reference consistency, and two speed tiers. 6 workflows. Run it in your browser.

Try Wan 3.0 Now → Browse All Models

Related Reading

Film and Animation Workflows on Floyo

AI Ad Creatives for Social and Web

Top AI Models on Floyo

Last updated: August 2026. Specs from Alibaba Wan 3.0 feature highlights document, Alibaba Cloud Model Studio pricing, QwenCloud wan3.0-video API reference (August 7, 2026), IT之家 and 第一财经 launch coverage (August 6, 2026), Pexo explainer, Siray Blog analysis, and Morphic model card.

TABLE OF CONTENTS
OVERVIEW

Run Wan 3.0 online through ComfyUI on Floyo. Alibaba's flagship video model with native 30-second single-shot generation, document-to-video (PDF/Excel/Word/PPT), web URL-to-video, native synchronized audio, pixel-level reference consistency, and smart duration control. Two tiers: Standard and Prime. 480p to 1080p. 6 workflows. No install, no GPU, browser-based. Free to try.