Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

Qwen3 for Voice Cloning

Clone any voice from a short audio clip using Qwen3 TTS, Alibaba's open-source speech model. Upload a reference, type your text, and hit run. Ten languages.

Audio
Audio to Audio
Voice Cloning

119

Gen time: ~51 secs

Nodes & Models

LoadAudio
Qwen3TTSEngineNode
UnifiedTTSTextNode
PreviewAudio

ABOUT THE WORKFLOW

Clone a Voice Upload a short audio clip of any voice, type the text you want spoken, and hear it read back in that voice. Supports ten languages. Works from as little as three seconds of reference audio.

Model

  • Qwen3 TTS 1.7B by Alibaba Cloud (Qwen team). Open-source text-to-speech model strong at zero-shot voice cloning, multilingual synthesis, and natural language emotion control. Apache 2.0 license.


HOW IT WORKS

Step 1. Upload a reference audio clip A short clip of the voice you want to clone. Three to ten seconds of clean, single-speaker audio works best. Works great with: voice recordings · podcast clips · dialogue lines

Step 2. Type the text Write the words the cloned voice will speak. The model supports Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian.

Step 3. Add a style instruction (optional) Describe the speaking style you want in a short sentence, like "speak softly and slowly" or "read with excitement." Leave it empty to let the model match the reference tone.

Step 4. Hit run and preview Qwen3 TTS generates the speech and plays it in the preview node on the canvas. Right-click the preview to download. Ready for: video editors · podcast tools · game engines · any DAW

First time? Leave every setting as-is. The defaults (1.7B model, auto language, 0.9 temperature) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard voice clone (most people) — 1.7B model, auto language, temperature 0.9, fixed seed. The right starting point for almost everyone.

  • Faster generation on lighter hardware — Switch the model size to 0.6B. Quality drops slightly, but generation is noticeably faster.

  • More expressive or varied delivery — Raise temperature toward 1.0 or higher. The voice becomes more dynamic but less predictable.

  • Consistent, controlled delivery — Lower temperature toward 0.5 to 0.7. The output stays closer to the reference tone with less variation.

  • Different take of the same text — Change the seed number. Each seed produces a different reading of the same text and reference.

  • Speaking in a specific style — Write a short instruction like "speak softly and slowly" or "narrate with a calm, warm tone" in the instruct field.

  • Long script or narration — The workflow chunks long text automatically (400 characters per chunk by default). For longer scripts, split into separate runs to avoid token limit cutoffs.

  • Voice sounds off or distorted — Use a cleaner reference clip. Background noise, music, or multiple speakers in the reference degrade the clone. Keep the clip under ten seconds.

Prompt: Write the full text the voice will speak. Punctuation affects delivery: commas add pauses, periods add stops, question marks shift intonation. For the instruct field, describe the mood or pace in one sentence. "Read with warmth and a slow pace" is clearer than "sound nice."


LEARN

📹 Videos

✨ Quick links


USE CASES

🎙️ Voiceover & Narration Clone a narrator's voice and generate lines for videos, explainers, or audiobooks without re-recording.

🎮 Game & Interactive Media Generate NPC dialogue in a consistent voice across scenes. Iterate on delivery by changing the seed or instruction.

🌍 Multilingual Content Record one reference clip and generate speech in any of the ten supported languages, keeping the same voice identity across translations.

🎬 Video Production Fill in placeholder dialogue, generate scratch tracks, or produce alternate reads for edit review without calling talent back to the booth.

🧪 Prototyping & Testing Test how voice lines sound in an app, product, or interface before committing to a full recording session.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clean, single-speaker reference clips (3 to 10 seconds)

  • Well-punctuated text with clear sentence structure

  • Short to medium text lengths per run

  • Style instructions that name a specific mood or pace

⚠️ May produce softer results

  • Noisy reference audio or clips with background music

  • Reference clips longer than 10 to 15 seconds

  • Multiple speakers in the reference clip

  • Long scripts that push the token limit in a single run


FAQ

What is Qwen3 TTS and who made it? Qwen3 TTS is an open-source text-to-speech model built by the Qwen team at Alibaba Cloud. It supports voice cloning from as little as three seconds of audio, covers ten languages, and is available in 0.6B and 1.7B parameter sizes. It is released under the Apache 2.0 license.

How many seconds of audio do I need to clone a voice? Three seconds of clean speech is enough. The model works best with three to ten seconds of a single speaker with no background noise. Longer clips can cause slower generation or artifacts, so keep the reference short and clean.

What languages does Qwen3 TTS support? Ten languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian. You can clone a voice in one language and generate speech in another.

Can I control how the cloned voice sounds beyond the reference? Yes. The instruct field accepts a short natural language description of the speaking style you want. Write something like "speak softly and slowly" or "narrate with energy and enthusiasm." The model adjusts tone, pace, and expression based on the instruction.

Is Qwen3 TTS free to use commercially? Yes. The model is open source under the Apache 2.0 license, so the weights are free to use for personal and commercial projects. On this workflow, you pay for generation time only.

How is Qwen3 TTS different from ElevenLabs or other voice cloning tools? Qwen3 TTS is fully open source and runs locally rather than through a proprietary API. You own the pipeline. It supports ten languages, cross-lingual cloning, and natural language style control. Commercial alternatives like ElevenLabs offer polished UIs and hosted APIs, but charge per character or per minute and keep the model closed.

How to run Qwen3 TTS online? You can run Qwen3 TTS online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload a reference audio clip, type your text, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Upload a voice clip, type what you want it to say, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N