Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

Fish Audio S2.1 for Voice Cloning

Clone a voice from short audio samples and generate speech from text using Fish Audio S2.1 Pro. Upload recordings, type your script, and hit run.

Audio to Audio
Fish Audio S2.1
Voice Cloning

51

Gen time: ~18 secs

Nodes & Models

FishAudioCreateVoiceModel_floyo
FishAudioTTSAdvanced_floyo
LoadAudio
PreviewAudio

ABOUT THE WORKFLOW

Clone a Voice and Generate Speech
Upload one to three short audio recordings of a voice. Type the script you want spoken. The workflow trains a voice model from your samples and reads your script back in that voice. You get a downloadable audio file.

Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.

Model

  • Fish Audio S2.1 Pro by Fish Audio. A production TTS model strong at voice cloning from short samples, multilingual speech in 83 languages, and inline emotion control with bracket tags like [whisper] or [excited].


HOW IT WORKS

Step 1. Upload your voice samples
Short audio clips of the voice you want to clone. 10 to 30 seconds each, one speaker per clip, minimal background noise. One clip is required. Two more are optional and improve the clone.
Works great with: clean dialogue · voiceover recordings · podcast clips

Step 2. Add transcripts (optional)
Type the exact words spoken in each audio clip. Helps accuracy, especially for non-English audio.

Step 3. Write your script
The text you want spoken in the cloned voice. Any language the model supports. Add emotion tags in brackets for expressive delivery, like "[excited] Great news!" or "[whisper] Listen closely."

Step 4. Hit run and preview
The workflow trains a voice model from your samples, then reads your script in the cloned voice. Preview the audio in the workflow, then download the file.
Ready for: video editing · podcasts · audiobooks · game engines · voice agents

First time? Leave every setting as-is. The defaults (fast training, WAV output, balanced latency) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Standard voice clone (most people) — 1 to 3 audio samples · fast train mode · WAV · temperature 0.7. The right starting point for almost everyone.

  • Higher quality clone — Upload all three audio sample slots with clips from different sentences. More variety in the samples gives the model a wider range of sounds to learn from.

  • Cloning a non-English voice — Add transcripts for each audio clip. The model uses them to match pronunciation more accurately in languages other than English.

  • Want more expressive delivery — Add bracket tags in your script like [excited], [whisper], or [laughing nervously]. The model reads them as style directions, not spoken text.

  • The clone sounds off or flat — Check your source audio. Background noise, echo, or overlapping speakers degrade the result. Re-record in a quiet room if possible.

  • Need consistent results across runs — Lower the temperature toward 0.3. The default (0.7) adds natural variation; lower values make each take sound closer to the last.

Prompt: Write your script in full sentences. Short fragments produce choppy output. "The quarterly results exceeded expectations, and the board approved the next phase" reads better than "Quarterly results. Board approved." For emotion, place bracket tags before the phrase they apply to.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎙️ Voiceover & Narration
Clone a narrator's voice and generate hours of consistent audio for videos, courses, or internal presentations without re-recording.

🌍 Multilingual Localization
Clone one voice and generate speech in multiple languages. The model supports 83 languages with automatic detection, so one speaker can cover localized content across markets.

🎮 Game & Interactive Audio
Generate character dialogue from a cloned voice. Use bracket emotion tags to shift tone between lines without recording separate takes.

📚 Audiobook Production
Turn a manuscript into spoken audio in a consistent voice. Break long scripts into sections and run them through the same voice model.

🤖 Voice Agents & Prototyping
Build a branded voice for a product demo or voice agent. Clone the voice once, then generate new lines from text as the script changes.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clean recordings of one speaker, 10 to 30 seconds each

  • Samples with varied sentences and natural pacing

  • Full-sentence scripts with punctuation

  • Bracket emotion tags for expressive control

⚠️ May produce softer results

  • Noisy recordings with background music or echo

  • Samples with multiple overlapping speakers

  • Very short fragments or single-word prompts

  • Extremely long scripts in a single pass (chunk into sections instead)


FAQ

What is Fish Audio S2.1 Pro and how does voice cloning work?
Fish Audio S2.1 Pro is a production text-to-speech model made by Fish Audio. It supports 83 languages, voice cloning, and inline emotion control. For cloning, you upload short audio samples (10 to 30 seconds), and the model builds a voice profile from the timbre, speaking style, and rhythm in those clips. It then reads any text you provide in that cloned voice, with no additional fine-tuning needed.

How long do my voice samples need to be for a good clone?
10 to 30 seconds per clip is the recommended range. One sample is the minimum. Adding a second and third clip with different sentences gives the model a wider range of sounds to learn from, which improves the result. The clips should be clean recordings of a single speaker with minimal background noise.

What languages does Fish Audio S2.1 support for voice cloning?
S2.1 Pro supports 83 languages, including English, Japanese, Chinese, Korean, Spanish, French, German, Arabic, and Russian. The model detects the language automatically from your text. For non-English cloning, adding a transcript of what is spoken in each audio sample improves pronunciation accuracy.

Can I control emotion and speaking style in the generated speech?
Yes. Fish Audio S2.1 supports natural-language bracket tags placed in your script. Tags like [whisper], [excited], [laughing nervously], or [professional broadcast tone] tell the model how to deliver the next phrase. These are not limited to a fixed set. You can write custom style directions in brackets.

Is the output from Fish Audio S2.1 safe for commercial use?
Fish Audio S2.1 Pro is a proprietary API model. Commercial use of the generated audio is allowed under Fish Audio's terms of service. Make sure you have the rights to clone any voice you upload, and review Fish Audio's current usage policy for your specific use case.

How is Fish Audio S2.1 different from ElevenLabs or other TTS tools?
Fish Audio clones a voice from a short sample without additional training time or per-voice fees. It supports 83 languages with automatic detection and offers open-ended emotion control through bracket tags rather than a fixed list of presets. The S2.1 Pro model also reports lower time-to-first-audio latency (around 70 ms) compared to the previous generation.

How to run Fish Audio S2.1 voice cloning online?
You can run Fish Audio S2.1 voice cloning online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload your voice samples, type your script, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Upload a voice sample, type your script, and run it. The settings are already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N