Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

CosyVoice3 for Voice Conversion

Convert any voice recording into a different speaker using CosyVoice 3 by Alibaba. Upload two audio clips, hit run, and get it re-spoken in a new voice.

Audio
Audio to Audio
Voice Change
Voice Conversion

76

Gen time: -- secs

Nodes & Models

LoadAudio
FL_CosyVoice3_ModelLoader
FL_CosyVoice3_VoiceConversion
PreviewAudio

ABOUT THE WORKFLOW

Convert a Voice Recording Upload a source recording and a target voice sample. CosyVoice 3 rebuilds the source speech using the target speaker's voice. You get back one audio clip with the original words spoken in the new voice.

Model

  • CosyVoice 3 (Fun-CosyVoice3-0.5B) by FunAudioLLM, Alibaba Tongyi Lab. A 0.5B-parameter voice model strong at zero-shot voice conversion and cloning across 9 languages and 18+ Chinese dialects. Apache 2.0 license.


HOW IT WORKS

Step 1. Upload your source audio The recording whose words you want to keep. The model preserves what is said and replaces the voice. Works great with: spoken dialogue · voiceover takes · narration clips

Step 2. Upload your target voice A sample of the voice you want the output to sound like. A clean clip with one speaker and no background noise works best. A few seconds of clear speech is enough.

Step 3. Hit run and preview CosyVoice 3 rebuilds the source speech in the target speaker's voice and plays it back. Right-click the preview to download the audio file. Ready for: Premiere Pro · DaVinci Resolve · Audacity · any audio editor

First time? Leave every setting as-is. The defaults (speed 1x · random seed) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard conversion (most people) — Speed 1 · random seed. The right starting point for almost everyone.

  • Want a slower or faster delivery — Set speed below 1 to slow down the output, or above 1 to speed it up. Useful when matching pacing to a video timeline.

  • Want to reproduce the same result — Lock the seed to a fixed number. Running the same inputs with the same seed returns the same output.

  • Want a few variations to choose from — Run multiple times with different random seeds and compare the takes.

  • The output sounds unnatural — Use a cleaner target reference. Background noise, music, or overlapping speakers in the target clip degrade the conversion.

  • The output sounds muffled or thin — Match recording quality between source and target. If one is a phone recording and the other is studio audio, the gap shows in the result.

Prompt: This workflow has no text prompt. The voice conversion is driven entirely by the two audio inputs. The source defines the words and timing. The target defines the voice.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎙️ Voiceover & Narration Re-record a voiceover in a different speaker's voice without re-hiring talent or booking a studio session.

🎬 Film & Post-Production Replace a scratch track with a matched voice for temp mixes, ADR previews, or multilingual dubs where you need the same delivery in a new voice.

🎓 Content Localization Convert training videos, course narration, or explainer audio into a consistent voice across your library.

🎵 Creative Audio Projects Experiment with vocal identity for music demos, podcast pilots, or character voice prototyping without recording new takes.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clean single-speaker recordings

  • Target clips with clear, noise-free speech

  • Matched recording quality between source and target

  • Short to medium source clips

⚠️ May produce softer results

  • Noisy or reverb-heavy target samples

  • Multiple overlapping speakers in either clip

  • Long single-pass source clips (quality can drift)

  • Large quality gap between source and target recordings


FAQ

What is CosyVoice 3 and who made it? CosyVoice 3 is a 0.5B-parameter speech model by FunAudioLLM, the speech team at Alibaba's Tongyi Lab. It handles text-to-speech, voice cloning, and voice conversion. This workflow uses the voice conversion mode, which takes two audio inputs and re-speaks one in the other's voice.

How does CosyVoice 3 voice conversion work? You upload two audio files. The source defines what is said and how it is timed. The target defines the voice. CosyVoice 3 reconstructs the source speech using the target speaker's vocal characteristics. No text prompt is involved. The output is 24 kHz mono audio.

What languages does CosyVoice 3 support for voice conversion? CosyVoice 3 covers 9 languages: Chinese, English, Japanese, Korean, German, Spanish, French, Italian, and Russian. It also supports 18+ Chinese dialects. Cross-lingual conversion works, so your source and target can be in different languages.

How long should the target voice sample be? A few seconds of clear speech is the minimum. Longer, cleaner samples give the model more to work with. The key factors are one speaker, no background noise, and consistent recording quality.

Is the output from CosyVoice 3 licensed for commercial use? Yes. CosyVoice 3 is released under the Apache 2.0 license, which permits commercial use. Outputs generated on this platform carry full commercial rights. Be sure you have the rights to use any voice you upload as a target.

What is the output quality and format? Output is 24 kHz mono audio. That is suitable for web, social media, voiceover, and most production workflows. It is not high-fidelity stereo, so plan on post-processing if your final delivery requires 48 kHz stereo or higher.

How to run CosyVoice 3 voice conversion online? You can run CosyVoice 3 voice conversion online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your two audio files, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A sound designer runs a conversion and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Upload your source recording and a target voice, then hit run.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N