Qwen3 TTS DVA for Voice Cloning
Upload a 3-second voice clip, type what you want it to say, and Qwen3-TTS speaks it in that voice. Ten languages, built-in emotion controls, open source.
Audio
Audio to Audio
Qwen3 TTS
Voice Cloning
63
Nodes & Models
QwenTTSModelLoader
LoadAudio
QwenTTSVoiceClone
QwenTTSEmotionMixer
PreviewAudio
ShowText|pysssss
ShowText|pysssss
ABOUT THE WORKFLOW
Clone a Voice from a Short Audio Clip Upload a short voice sample and type what you want it to say. Qwen3-TTS learns the voice from as little as three seconds, then speaks your text in that voice. Whisper transcribes the reference audio automatically, so you only need to supply the clip and your new text.
Model
Qwen3-TTS 1.7B Base by the Qwen team at Alibaba Cloud. An open-source voice cloning model that reproduces a speaker's voice from a 3-second sample across 10 languages, with built-in emotion controls. Licensed under Apache 2.0.
HOW IT WORKS
Step 1. Upload a voice sample A short audio clip of the voice you want to clone. Three seconds is enough. Longer clips work but are not required. Works great with: clean recordings · single speaker · minimal background noise
Step 2. Type the text you want spoken Write the words the cloned voice should say. Can be in any of the 10 supported languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, or Italian.
Step 3. Hit run and preview Whisper transcribes your reference clip, the voice model learns the speaker, and the emotion mixer shapes the final output. You hear the result in the preview player. Ready for: podcasts · voiceovers · dubbing · prototyping
First time? Leave every setting as-is. The defaults (Auto language · low temperature · random seed) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard voice clone (most people) — Auto language · temperature 0.1 · random seed · default emotion values. The right starting point for almost everyone.
More expressive delivery — Raise temperature toward 0.2 or 0.3. Higher values add variation to intonation. Stay below 0.5 to keep speech stable.
Reproduce a specific result — Lock the seed to a fixed number. This gives the same output each time so you can adjust one setting at a time.
Adjust emotion and delivery — Use the emotion controls: tempo, pitch, energy, brightness, warmth, and articulation. The defaults sit near 0.5. Small changes go a long way. Extreme values distort the output.
Clone across languages — Upload a reference in one language, type text in another. Qwen3-TTS handles cross-lingual synthesis, though matching the reference language to the output language gives the closest voice match.
The voice sounds off — Try a cleaner reference clip. Background noise, music, or overlapping speakers reduce accuracy. A clear, single-speaker recording is the single biggest quality factor.
Prompt: Type the words you want spoken. Keep sentences at a natural reading length. Long passages can slow generation and lose consistency, so break scripts into shorter blocks and run them one at a time.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎙️ Voiceover Production Clone a narrator's voice and generate scratch tracks or alternate takes without rebooking studio time.
🌍 Multilingual Dubbing Take a speaker's voice in one language and have it speak in another. Useful for localizing video content, tutorials, or presentations across markets.
🎧 Podcast and Audio Prototyping Test how a script sounds in a specific voice before committing to a full recording session.
🎮 Game and Animation Dialogue Generate placeholder or final character dialogue from a short voice sample. Iterate on delivery by adjusting emotion controls between runs.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Clean recordings with a single speaker
Three seconds or more of clear speech
Matching the reference language to the output language
Short to medium text passages
⚠️ May produce softer results
Noisy reference audio or background music
Multiple overlapping speakers in the sample
Long text passages in a single run
Extreme emotion values pushed beyond the stable range
FAQ
What is Qwen3-TTS and who made it? Qwen3-TTS is an open-source text-to-speech model family developed by the Qwen team at Alibaba Cloud, released in January 2026. The 1.7B Base variant used in this workflow specializes in voice cloning. It takes a short audio sample and text, then generates speech in the cloned voice. Licensed under Apache 2.0 for research and commercial use.
How long does the voice sample need to be for Qwen3-TTS voice cloning? Three seconds is enough. The model extracts the speaker's voice characteristics from that short reference. Longer clips give the model more to work with, but the gains diminish quickly. A clean three-second clip of one person speaking clearly outperforms a noisy thirty-second recording.
What languages does Qwen3-TTS support for voice cloning? Qwen3-TTS supports 10 languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian. You can clone a voice in one language and have it speak in another, though the closest voice match comes from keeping the reference and output in the same language.
Can I use Qwen3-TTS voice clones for commercial projects? Yes. The model is released under Apache 2.0, which permits commercial use. You are responsible for having the right to use the voice you clone and for complying with applicable laws around synthetic speech and consent in your jurisdiction.
What do the emotion controls do in Qwen3-TTS? The emotion mixer adjusts six properties of the generated speech: tempo, pitch, energy, brightness, warmth, and articulation. Defaults sit near 0.5 on each axis. Small adjustments shape delivery. For example, raising energy and tempo produces a more upbeat read. Keep values moderate. Extreme settings distort the audio.
Is Qwen3-TTS better than ElevenLabs or other commercial voice cloning tools? Qwen3-TTS is open-source and runs locally or on cloud GPUs, so you control your data and pay only for compute. Commercial services like ElevenLabs offer polished APIs and additional post-processing. For teams that need data privacy, Apache 2.0 licensing, and control over the generation pipeline, Qwen3-TTS is a strong option.
How to run Qwen3-TTS voice cloning online? You can run Qwen3-TTS voice cloning online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload a voice sample, type your text, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A sound designer clones a voice and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it? Upload a voice clip, type what you want it to say, and run it. The settings are already set.
Questions? Watch the free course or check the FAQ above.
Read more







