Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

IndexTTS 2.5 and TTS Audio Suite for Voice Cloning

Clone any voice and shape its emotion with IndexTTS-2. Upload a short reference clip, type your text, set the mood with eight sliders, and run.

Audio
IndexTTS 2.5
Voice Clone

82

Gen time: ~3 min 32 secs

Nodes & Models

IndexTTSEmotionOptionsNode
MarkdownNote
LoadAudio
IndexTTSEngineNode
UnifiedTTSTextNode
PreviewAudio

ABOUT THE WORKFLOW

Clone a Voice with Emotion Upload a short audio clip of any voice, type the text you want spoken, and adjust eight emotion sliders to shape the delivery. IndexTTS-2 clones the voice from your reference and speaks your text with the emotion you set.

Model

  • IndexTTS-2 by Bilibili. A zero-shot text-to-speech model that clones a voice from a short reference clip and generates speech with independent emotion control across eight dimensions.


HOW IT WORKS

Step 1. Upload a voice reference A short audio clip of the voice you want cloned. Cleaner recordings with minimal background noise produce the closest match. Works great with: voice recordings · podcast clips · audiobook samples

Step 2. Type your text Enter the text you want spoken in the cloned voice.

Step 3. Set the emotion Eight sliders control the emotional tone: Happy, Angry, Sad, Surprised, Afraid, Disgusted, Calm, and Melancholic. Each runs from 0.0 to 1.2. Mix them to shape the delivery.

Step 4. Hit run and preview The workflow generates the audio and shows a preview player. Listen, then download. Ready for: video editors · DAWs · podcasts · any audio tool

First time? Leave every setting as-is. The defaults produce a natural, lightly emotional read that works for most use cases.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard voice clone (most people) — Default emotion sliders, seed fixed, chunking on. The right starting point for almost everyone.

  • Want a specific emotional tone — Push one emotion slider to 0.6 or higher and keep the rest low. For example, set Happy to 0.8 and everything else below 0.2 for an upbeat read.

  • Want a complex mixed emotion — Combine two or three sliders at moderate values. Happy at 0.5 plus Surprised at 0.4 produces an excited tone. Start subtle and increase from there.

  • Need a longer script read smoothly — Keep chunking enabled (it is by default). The workflow splits long text into segments and stitches the audio together with a short silence gap between them.

  • Want variation between takes — Change the seed number. Each seed produces a different reading of the same text and emotion settings.

  • The cloned voice sounds off — Use a cleaner reference clip. Background music, reverb, and overlapping speakers all degrade the cloning accuracy. A 5 to 15 second clip of a single speaker in a quiet room works best.

  • Emotion is too intense or distorted — Lower the emotion alpha. The default is 0.7. Dropping it to 0.4 or 0.5 dials back the intensity while keeping the emotional direction.

Prompt: This workflow does not use a creative prompt. You type the exact words you want spoken. Punctuation matters: periods create pauses, commas create shorter breaks, and question marks shift the intonation upward.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Video Dubbing Clone a speaker's voice and generate dialogue lines that match the original tone. Use the emotion sliders to match the mood of each scene.

📖 Audiobook Narration Turn a manuscript into narrated audio with a consistent cloned voice. Adjust the emotion per chapter or scene for expressive readings without re-recording.

🎙️ Podcast Production Generate voice segments in a cloned voice for intros, transitions, or multilingual versions of existing episodes.

🎮 Game & App Voice Lines Produce character voice lines at scale. Set each line's emotion independently to match dialogue trees and story beats.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clean, single-speaker reference clips (5 to 15 seconds)

  • Short to medium text passages with clear punctuation

  • Single or paired emotion sliders at moderate values

  • Chinese and English text

⚠️ May produce softer results

  • Noisy reference audio with music or reverb

  • Emotion sliders pushed above 1.0 (may distort the cloned voice)

  • Long unbroken text without punctuation

  • Reference clips with multiple overlapping speakers


FAQ

What is IndexTTS-2 and who made it? IndexTTS-2 is a zero-shot text-to-speech model built by Bilibili's Index Team. It clones a voice from a short reference clip and generates speech with independent emotion and speaker control. "Zero-shot" means no fine-tuning is needed. You supply one audio reference and the model clones the voice on the spot.

How does the emotion control work in IndexTTS-2? Eight sliders represent different emotions: Happy, Angry, Sad, Surprised, Afraid, Disgusted, Calm, and Melancholic. Each slider runs from 0.0 (none) to 1.2 (intense). You can mix emotions by raising multiple sliders at once. The model separates emotion from speaker identity, so changing the mood does not change whose voice it sounds like.

What languages does IndexTTS-2 support? IndexTTS-2 supports Chinese and English. It can clone a voice from one language and generate speech in the other (cross-lingual synthesis), though results are strongest when the reference and output language match.

How long should the reference audio clip be? Five to fifteen seconds of clear, single-speaker audio works best. The clip should have minimal background noise, no music, and no overlapping voices. Longer clips are fine but do not improve quality much past 15 seconds. Shorter clips below 3 seconds may produce less accurate cloning.

Is IndexTTS-2 free for commercial use? The code is released under Apache 2.0. The model weights carry a separate Bilibili license that requires written authorization for commercial use. For non-commercial and personal projects, no additional permission is needed. For commercial projects, contact the Index Team at indexspeech@bilibili.com before using outputs in a shipped product.

What is the difference between IndexTTS-2 and other TTS models like Coqui XTTS or Bark? IndexTTS-2 separates emotion from speaker identity, so you can change the emotional tone without affecting the cloned voice. Most other zero-shot TTS models treat emotion and speaker as linked. IndexTTS-2 also supports precise duration control for video dubbing where timing matters. The tradeoff is that it runs heavier than lightweight models like Bark.

How to run IndexTTS-2 voice cloning online? You can run IndexTTS-2 voice cloning online through Floyo. No installation, no setup, no dependency management. Open the workflow in your browser, upload your reference audio, type your text, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A sound designer clones a voice and dials in the emotion. A teammate opens that exact run from shared history and generates the next batch of lines. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it? Upload a voice clip, type your text, and run it. The emotion sliders are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N