Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

CosyVoice 3 for Voice Clone

Clone any voice from a short audio clip using CosyVoice 3, Alibaba's multilingual zero-shot TTS model. Upload a voice sample, type your text, and hit run.

Audio
CosyVoice3
Voice Clone

69

Gen time: -- secs

Nodes & Models

LoadAudio
FL_CosyVoice3_ModelLoader
FL_CosyVoice3_ZeroShot
PreviewAudio

ABOUT THE WORKFLOW

Clone a Voice from a Short Clip Upload a few seconds of someone speaking and type the text you want spoken in that voice. CosyVoice 3 reads the reference clip and generates new speech that matches it. No training, no fine-tuning. One run, one result.

Model

  • CosyVoice 3 (0.5B) by FunAudioLLM at Alibaba Tongyi Lab. A zero-shot multilingual TTS model strong at speaker similarity and natural prosody across nine languages.


HOW IT WORKS

Step 1. Upload a voice sample A short recording of the voice you want to clone. Five to fifteen seconds of clear speech works best. Avoid clips with background music or noise. Works great with: voice recordings · podcast clips · dialogue samples

Step 2. Type your text Write the words you want spoken in the cloned voice. Supports Chinese, English, Japanese, Korean, German, Spanish, French, Italian, and Russian.

Step 3. Hit run and preview CosyVoice 3 generates speech matching the reference voice. The result plays back in the preview player. Right-click the audio to download it. Ready for: video editors · podcast tools · game engines · any audio pipeline

First time? Leave every setting as-is. The defaults are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard voice clone (most people) — Speed 1 · random seed · text frontend on. The right starting point for almost everyone.

  • Need a slower, more deliberate read — Lower the speed to 0.8 or 0.7. Good for narration and audiobook-style delivery.

  • Need a faster read for dialogue — Raise the speed to 1.2 or 1.3. Keeps the voice natural at a quicker tempo.

  • Want to reproduce the same result — Lock the seed to a specific number. Same seed plus same inputs gives the same output every time.

  • Want to compare variations — Leave the seed on random and run it a few times. Pick the take that sounds best.

  • The voice sounds off or robotic — Try a cleaner reference clip before changing settings. Background noise and music in the sample are the most common cause.

  • Working in a non-English language — Type the text in that language directly. CosyVoice 3 handles multilingual synthesis natively across nine languages.

Prompt: This workflow does not use a text prompt for style control. The text field is what gets spoken. Write it as you want it read aloud. Punctuation affects pacing: commas add pauses, periods add stops.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎙️ Voiceover & Narration Clone a narrator's voice and generate new lines without booking another recording session. Good for video essays, explainers, and internal training material.

🎮 Game & Animation Dialogue Prototype character voices from a short sample. Generate placeholder dialogue during production and iterate on delivery before the final recording session.

🌍 Multilingual Content Record a voice sample in one language and generate speech in another. CosyVoice 3 supports nine languages, so you can produce localized audio from a single reference.

🎧 Podcast & Audio Production Fill in missing lines, re-record a flubbed sentence, or extend a read without calling the speaker back. The clone matches the tone and cadence of the original.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clean speech recordings with minimal background noise

  • Reference clips between 5 and 15 seconds

  • Text in any of the nine supported languages

  • Natural, conversational speaking samples

⚠️ May produce softer results

  • Noisy recordings with music or ambient sound

  • Reference clips under 2 seconds

  • Languages outside the supported nine

  • Long text passages (voice may drift over time)


FAQ

What is CosyVoice 3 and who made it? CosyVoice 3 (Fun-CosyVoice 3.0) is a zero-shot text-to-speech model built by FunAudioLLM, the speech research team at Alibaba Tongyi Lab. It clones a voice from a short audio clip and generates new speech in that voice, with no fine-tuning required. The model is 0.5 billion parameters, built on a Qwen2.5 backbone with a flow-matching vocoder.

What languages does CosyVoice 3 support? CosyVoice 3 supports nine languages: Chinese, English, Japanese, Korean, German, Spanish, French, Italian, and Russian. It also handles 18+ Chinese regional dialects, including Cantonese, Sichuan, and Shanghainese. Cross-lingual cloning is supported, meaning you can provide a reference clip in one language and generate speech in another.

How long does the reference audio clip need to be? Five to fifteen seconds of clear, natural speech works best. Shorter clips give the model less to match, which can reduce similarity. Longer clips are fine, but the extra length does not improve results much beyond 15 seconds. The most important factor is audio quality: clean speech with no background noise or music.

Is CosyVoice 3 open source, and can I use the output commercially? The model weights are released under the Apache 2.0 license, which permits commercial use, modification, and redistribution. The model card includes a note that demo content is for academic purposes, but that applies to the provided examples, not the license itself. You own the audio you generate.

How is CosyVoice 3 different from other voice cloning models like Seed-VC or RVC? CosyVoice 3 is a text-to-speech model. You type text and it generates speech in a cloned voice. Seed-VC and RVC are voice conversion models: they take existing audio and change the voice while keeping the original timing and delivery. Use CosyVoice 3 when you need to generate new speech from text. Use voice conversion when you already have the audio and want to swap the speaker.

Does CosyVoice 3 produce natural-sounding speech? CosyVoice 3 is designed for natural prosody and speaker similarity from short clips. It handles pauses, intonation, and rhythm well for most content. Long passages may drift from the reference voice over time. For longer reads, splitting the text into shorter segments and running them separately gives more consistent results.

How to run CosyVoice 3 voice cloning online? You can run CosyVoice 3 voice cloning online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload a voice sample, type your text, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A sound designer runs a voice clone and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Upload a voice sample, type your text, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N