Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

VibeVoice · Text to Speech

Clone any voice from a short clip and read your text in it using VibeVoice Large by Microsoft. Upload a voice sample, type your script, hit run. MIT license.

1.3k

Gen time: ~34 secs

Nodes & Models

LoadTextFromFileNode
VibeVoiceSingleSpeakerNode
LoadAudio
PreviewAudio

ABOUT THE WORKFLOW

Read Text in a Cloned Voice Upload a short audio clip of the voice you want, type the text you want spoken, and hit run. VibeVoice clones the voice from your sample and reads your script in it. The result plays back in the workflow as a preview.

Model

  • VibeVoice Large by Microsoft. Open-sourced August 2025 under the MIT License. A text-to-speech model with voice cloning that reads from a short reference sample and generates speech matching its tone, pacing, and character. Built on a next-token diffusion framework with continuous speech tokenizers at 7.5 Hz. Supports English and Mandarin Chinese.


HOW IT WORKS

Step 1. Upload a voice sample A short audio clip of the voice you want cloned, around 10 seconds. Clear speech with minimal background noise produces the closest match. Works great with: voice memos · interview clips · podcast samples · narration recordings

Step 2. Type your script Write the text you want spoken into the node. Punctuation controls pacing. Line breaks add pauses.

Step 3. Hit run and listen The model clones the voice from your sample and reads your script in it. The result plays in the audio preview.

Step 4. Download Right-click the audio preview to save the file. There is no automatic save step. Ready for: Premiere · DaVinci Resolve · Audacity · any audio editor

First time? Leave every setting as-is. The defaults (VibeVoice-Large · 20 diffusion steps · CFG 1.3 · fixed seed) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard voice clone (most people) — VibeVoice-Large · 20 diffusion steps · CFG 1.3 · fixed seed. The right starting point for almost everyone.

  • Voice sounds robotic or flat — Try a cleaner reference clip. Background noise, music, or multiple speakers in the sample dilute the clone.

  • Voice drifts from the reference — Raise CFG toward 2. Higher values hold the output closer to the reference voice, but push too far and the speech sounds forced.

  • Pacing is too fast or too slow — Adjust voice speed. Lower than 1 slows it down, higher speeds it up.

  • Long text has audible seams — Text over 250 words is split into chunks. Seams between chunks can be audible. Shorten the text or place natural pauses (periods, line breaks) where the split might land.

  • Smoother output — Raise diffusion steps above 20. More passes refine the audio at the cost of generation time.

  • Want a different read — Change the seed number. The default runs on fixed, so the same sample, text, and seed return the same output.

  • Text from a file — A LoadTextFromFile node sits muted by default. Enable it and point it at a text file instead of typing into the node.

Script: Punctuation is your pacing tool. Periods create full stops. Commas create short pauses. Line breaks add longer pauses. "Hello there. (pause) Let's see how this sounds." reads differently from "Hello there, let's see how this sounds." Write the script the way you want it heard.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎙️ Narration and Voiceover Clone a narrator's voice from a short sample and generate reads for explainer videos, courses, or audiobooks.

🎧 Podcast and Dialogue Read scripted dialogue in a specific voice for podcast pilots, demo episodes, or prototyping.

📢 Ads and Product Audio Generate a branded voice read for a product spot without booking a studio session.

🌍 Multilingual Content Read the same script in English and Mandarin Chinese in the same cloned voice for bilingual content.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clean reference clips of 10 to 30 seconds with one speaker

  • Scripts with deliberate punctuation and line breaks for pacing

  • English and Mandarin Chinese text

  • Narration, voiceover, and dialogue reads

⚠️ May produce softer results

  • Reference clips with background music, noise, or multiple speakers

  • Scripts longer than 250 words in a single run (chunk seams appear)

  • Languages other than English and Mandarin

  • Expecting pixel-perfect voice matching on extreme vocal ranges


FAQ

What is VibeVoice? VibeVoice is a family of open-source voice AI models from Microsoft. VibeVoice Large, used in this workflow, is a text-to-speech model with voice cloning that reads from a short reference audio sample and generates speech matching its tone, pacing, and character. It was open-sourced in August 2025 and accepted as an Oral at ICLR 2026. It uses continuous speech tokenizers at 7.5 Hz and a next-token diffusion framework.

How long does the voice sample need to be? About 10 seconds of clear speech from a single speaker. Longer samples give the model more to work with, but diminishing returns set in past 30 seconds. Quality matters more than length: a clean 10-second clip outperforms a noisy 60-second one.

What languages does VibeVoice support? English and Mandarin Chinese, with cross-lingual generation that allows natural language switching within the same output. Other languages are not officially supported and produce lower quality results.

Is VibeVoice free for commercial use? The weights were released under the MIT License, which permits commercial use, modification, and redistribution. Microsoft's project page adds that VibeVoice is "intended for research and development purposes" and recommends further testing for commercial applications. The MIT licence is permissive, but read the responsible use guidance before deploying in production.

Why is there no save step? The workflow outputs to an audio preview node with no SaveAudio node wired after it. Right-click the preview to download the file. Adding a save node would make downloads automatic.

How long can the output be? VibeVoice Large can generate extended speech, but this workflow splits text into chunks of 250 words. Seams between chunks can be audible. For the cleanest output, keep each run under 250 words and join the clips in an editor.

How to run VibeVoice online? You can run VibeVoice online through Floyo. No installation, no setup, no model downloads. Open the workflow in your browser, upload a voice sample, type your script, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Upload a voice sample, type your script, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N