Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

Timbre Transfer (SeedVC) for Voice Change

Convert any voice into another with Seed-VC. Upload a source audio and a short reference clip, and the model rebuilds the audio in the reference voice.

Audio
Audio to Audio
SeedVC
Timbre Transfer
Voice Change
Voice Conversion

220

Gen time: ~1 min 2 secs

Nodes & Models

LoadAudio
PreviewAudio
SaveAudio

ABOUT THE WORKFLOW

Convert Any Voice Into Another Upload a source audio file and a short reference voice clip. The model extracts the content and melody from your source, copies the timbre from the reference, and rebuilds the audio in the new voice. Works for both speech and singing.

Model

  • Seed-VC (44k) by Plachtaa (Liu Songting). A zero-shot voice conversion model tuned for singing at 44.1 kHz. No training on the target speaker needed. Open weights under GPL-3.0.


HOW IT WORKS

Step 1. Upload your source audio The speech or song you want to convert. The words, melody, and rhythm carry over from this file. Works great with: vocals · songs · speech recordings · voiceovers

Step 2. Upload a reference voice clip A short sample of the voice you want the output to sound like. 6 to 20 seconds of clean, single-speaker audio works well.

Step 3. Hit run The model swaps the voice while keeping the original lyrics, timing, and pitch. A stereo restoration step rebuilds the spatial width from your original source. Preview and download the result. Ready for: DAWs · video editors · podcasts · music production

First time? Leave every setting as-is. The defaults (60 diffusion steps, full timbre strength, no pitch shift) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard voice swap (most people) — 60 steps, timbre strength 1, pitch shift 0, auto F0 off. The right starting point for almost everyone.

  • Quick preview or test — Drop diffusion steps to 10. Faster, lower quality, but good enough to check if the pairing works.

  • Best quality for a final render — Raise diffusion steps to 80 or 100. Cleaner conversion, slower generation.

  • Cross-gender conversion — Turn auto F0 adjust on. This matches the pitch range to the reference voice so male-to-female or female-to-male swaps sound natural.

  • Subtle voice blend instead of full swap — Lower timbre strength to 0.5 or 0.6. The output keeps more of the original voice character.

  • Pitch correction after conversion — Use the pitch shift slider to nudge the output up or down in semitones. Useful when the reference sits in a different register than the source.

  • Reference clip is noisy or has reverb — Record or find a cleaner sample. Background noise in the reference bleeds into the output timbre.

Prompt: No text prompt needed. This workflow is audio-in, audio-out. The model reads content from your source file and timbre from your reference clip.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎤 Cover Artists & Musicians Sing a track in your own voice, then convert it to match a different vocal style or register. Preview how a song sounds in another voice before committing to a session singer.

🎬 Filmmakers & Post-Production Replace a scratch vocal with a different voice while preserving the original performance timing. Useful for dubbing, ADR placeholders, and voice matching across takes.

🎙️ Podcast & Voiceover Producers Swap a narrator voice across episodes for consistency, or test how a script sounds in a different vocal character before recording the final take.

🎵 Music Producers Convert demo vocals into a different timbre to explore arrangement ideas. The 44 kHz output and stereo restoration keep the result production-ready.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clean, isolated vocals or speech

  • Reference clips of 6 to 20 seconds with one speaker

  • Singing voice conversion (this 44k checkpoint is tuned for it)

  • Cross-gender swaps with auto F0 adjust turned on

⚠️ May produce softer results

  • Noisy or reverb-heavy reference clips

  • Source audio longer than 30 seconds (gets chunked, may have seams)

  • Reference clips under 3 seconds

  • Source recordings with heavy background music or multiple speakers


FAQ

What is Seed-VC and how does it work? Seed-VC is a zero-shot voice conversion model by Plachtaa. It takes two audio inputs: a source (the words and melody to keep) and a reference (the voice to copy). It extracts content from the source and timbre from the reference, then rebuilds the audio in the new voice. No training on the target speaker is required.

Is Seed-VC good for singing voice conversion? Yes. The 44k checkpoint used in this workflow is tuned for singing. It preserves lyrics, melody, and rhythm through the conversion while swapping the vocal timbre. In evaluations, it outperformed speaker-specific RVC models in both speaker similarity and intelligibility, despite being zero-shot.

How long should the reference voice clip be? Between 6 and 20 seconds of clean, single-speaker audio. Longer is not always better. The model auto-trims clips beyond 25 seconds. Focus on a segment with clear speech or singing, minimal background noise, and no overlapping voices.

Can Seed-VC do cross-gender voice conversion? Yes. Turn on the auto F0 adjust setting. This matches the pitch range of the output to the reference voice, so male-to-female and female-to-male conversions sound natural instead of pitched-up or pitched-down.

What is the difference between Seed-VC and RVC? RVC models are trained on a specific speaker and work best for that one voice. Seed-VC is zero-shot: it clones any voice from a short reference clip with no training. RVC can produce higher fidelity for a single trained speaker, but Seed-VC handles any voice on the fly.

Is Seed-VC open source and can I use it commercially? Seed-VC is open source under the GPL-3.0 license. The model weights and code are freely available. GPL-3.0 requires derivative works to remain open source under the same license. Check the license terms if you plan to distribute software that incorporates the model.

How to run Seed-VC voice conversion online? You can run Seed-VC voice conversion online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your audio files, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A producer runs a voice conversion and likes the result. A bandmate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Upload a source audio, drop in a reference voice clip, and hit run.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N