Gemini 3.8 Flash TTS for Text to Speech
Turn any script into spoken audio using Google's Gemini 3.8 Flash TTS. Pick a voice, describe the delivery style, and hit run.
Audio
Gemini 3.8 Flash
Text to Speech
TTS
47
ABOUT THE WORKFLOW
Turn Text Into Speech
Type or paste your script, pick a voice, and describe how it should sound. The model reads your text aloud and returns an audio file. Add speaker tags for multi-voice dialogue.
Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.
Model
Gemini 3.8 Flash TTS by Google. Strong at expressive narration, character voice design, multi-speaker dialogue, and accent control. Supports 130+ languages with automatic detection.
HOW IT WORKS
Step 1. Write or paste your script
The text the model reads aloud. A single line, a paragraph, or a full page.
Step 2. Pick a voice
Each voice name is a different speaker character. The default (Achernar) works for most narration.
Step 3. Describe the delivery style (optional)
Write a plain-language note on tone, pace, emotion, and accent. Example: "warm and calm, slow pace, British accent, bedtime story narrator."
Step 4. Set up multi-speaker dialogue (optional)
For conversations, tag each speaker in your script and assign a voice to Speaker 1 and Speaker 2.
Step 5. Hit run and preview
The model reads your script and returns a WAV audio file. Preview it in the built-in player, then download.
Ready for: Premiere Pro · DaVinci Resolve · Audacity · any audio or video editor
First time? Leave every setting as-is. The defaults (Achernar voice, preset style instructions) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard narration (most people) — Default voice (Achernar) · default style instructions. The right starting point for almost everyone.
Want a different tone or pace — Rewrite the style instructions before switching voices. "Excited and fast, podcast host energy" gives a different result from the same voice.
Need a specific accent — Add the accent to the style instructions. Example: "neutral Australian accent, documentary narrator."
Recording a conversation or interview — Tag speakers in your script, then assign a voice to Speaker 1 and Speaker 2 in the settings.
Working in another language — Paste your script in that language. The model detects it automatically. No language setting to change.
The delivery sounds flat — Rewrite the style instructions with more specific emotion and pacing cues before trying a different voice.
Prompt: Write your script as you want it read. Use punctuation to control pacing. For style instructions, be specific: "warm and enthusiastic, steady mid-pace, approachable tech narrator" lands better than "make it sound good."
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎙️ Video Narration
Generate voiceovers for YouTube videos, tutorials, or product demos without booking a voice actor or setting up a mic.
🎧 Podcast & Audio Content
Turn scripts into spoken audio for podcast intros, audio articles, or internal team updates.
🎭 Multi-Speaker Dialogue
Produce conversations, interviews, or character dialogue with two distinct voices in a single run.
🌍 Multilingual Voiceovers
Record narration in 130+ languages from the same workflow. Paste the script in the target language and run.
🎮 Game & App Prototyping
Generate placeholder voice lines for game characters, app onboarding flows, or interactive prototypes.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Clear, well-punctuated scripts
Specific style instructions (tone, pace, accent, emotion)
Short to mid-length narration
Two-speaker dialogue with tagged speakers
⚠️ May produce softer results
Very long scripts in a single pass
Vague style instructions like "make it sound nice"
Scripts with no punctuation or formatting cues
More than two speakers in one run
FAQ
What is Gemini 3.8 Flash TTS?
Gemini 3.8 Flash TTS is Google's text-to-speech model, released in September 2026. It takes a script, a voice selection, and optional style instructions describing tone, pace, emotion, and accent, and outputs a WAV audio file. It supports 130+ languages with automatic language detection.
How do style instructions work in Gemini 3.8 Flash TTS?
Style instructions are a plain-language description of how the voice should deliver the text. You write what you want in natural language, like "calm and measured, slow pace, deep voice, nature documentary narrator." The model follows these cues when reading your script. If the result does not match, rewrite the style instructions with more specific detail before switching to a different voice.
Can Gemini 3.8 Flash TTS do multiple voices in one audio file?
Yes. Tag each speaker in your script text, then assign a voice to Speaker 1 and Speaker 2 in the workflow settings. The model reads each part in the assigned voice. This works for dialogues, interviews, and two-character scenes.
What languages does Gemini 3.8 Flash TTS support?
It supports over 130 languages with automatic language detection. Paste your script in the target language and the model detects it on its own. No language selector to configure.
Does Gemini 3.8 Flash TTS support voice cloning?
No. This workflow does not include voice cloning. You select from a set of built-in voice characters and control the delivery through style instructions. Each voice name (Achernar, Charon, Kore, and others) is a distinct speaker with its own character.
Can I use the audio output commercially?
Gemini 3.8 Flash TTS is a proprietary Google model, so commercial use is governed by Google's terms. Commercial use is generally allowed, but review Google's current terms for your specific use case.
How to run Gemini 3.8 Flash TTS online?
You can run Gemini 3.8 Flash TTS online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, paste your script, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it?
Paste your script, pick a voice, and run it. The style instructions are already set.
Questions? Watch the free course or check the FAQ above.
Read more






