Floyo
Floyo
Workflows
API
Pricing
Floyo
Floyo
Workflows
API
Pricing

Gemini 3.8 Flash Lite TTS for Text to Speech

Turn any script into spoken audio using Google's Gemini 3.8 Flash-Lite TTS. Pick a voice, set the tone, and hit run.

Audio
Gemini 3.8 Flash Lite
Text to Speech
TTS

32

Gen time: -- secs

Nodes & Models

Gemini38FlashLiteTTS_floyo
PreviewAudio

ABOUT THE WORKFLOW

Turn Text Into Spoken Audio
Type a script, pick a voice, and get audio back. Add style instructions to shape the delivery. For dialogue, assign two speakers and the model reads both parts in one run.

Partner node. This workflow calls an external API, so each run uses credits from your API wallet. No API key needed. Floyo handles the connection.

Model

  • Gemini 3.8 Flash-Lite TTS by Google. The fast, cost-efficient member of the 3.8 TTS family. Strong at high-volume production work: dubbing, voice agents, audio content at scale. Ranked second on the Hume AI Voice Quality Index.


HOW IT WORKS

Step 1. Write your script
Type exactly what you want spoken. Use <short pause> and <long pause> tags inline to control timing. Add <laugh>, <sigh>, <gasp>, or <chuckle> where you want vocal sounds. Scripts can run up to 8,192 tokens, enough for a full scene.

Step 2. Pick a voice
Choose from 30 built-in voices. Kore is firm, Puck upbeat, Charon informative, Zephyr bright, Enceladus breathy, Sulafat warm, plus 24 more.

Step 3. Set the style (optional)
Add a short phrase describing how the voice should sound: "warm and enthusiastic," "whispered urgently," or "out of breath." Short phrases land better than sentences.

Step 4. Set up dialogue (optional)
Fill in both speaker IDs and pick a voice for each. Write your script as Name: what they say. The model reads both parts in a single run with no stitching needed.

Step 5. Hit run and download
The model generates one audio clip. Preview it in the player, then right-click to download.
Ready for: Premiere Pro · DaVinci Resolve · Audacity · any audio editor

First time? Leave every setting as-is. The defaults (Algenib voice, single speaker) are the right starting point for almost everyone.


RECOMMENDED SETTINGS

Quick-start guide. Find the goal that matches yours and copy the settings.

  • Single voice narration (most people) — Pick a voice, leave both speaker ID fields empty, hit run. That is all.

  • Two-speaker dialogue — Fill in both speaker IDs (e.g. "Maya" and "Arjun"), pick a voice for each. Write every line as Name: what they say. A line that names no speaker joins the line above it.

  • Want a specific tone or mood — Add a short phrase in style instructions: "deep gravitas, slow cinematic pacing" or "upbeat, sports commentary." This shapes the whole read.

  • Want one line to break the mood — Put the style in square brackets before the colon: Arjun [whispered urgently]: Keep this between us. Lines without brackets fall back to the style instructions.

  • Adding vocal sounds — Place tags inline in the text: <laugh> <sigh> <gasp> <breath> <chuckle> <cough> <yawn> <whispers>. Keep tags in English even when the surrounding text is not. No sound effects, music, or applause. Only sounds a person makes.

  • The voice is reading your instructions out loud — The model reads your prompt word for word. It does not follow instructions written inside the script. Delivery goes in style instructions or per-line brackets, never in the prompt itself.

Prompt: Write the script exactly as you want it spoken. "In a world drowning in noise <short pause> one image can change everything" is better than "Say this in an epic voice: one image can change everything." Delivery instructions go in style instructions, not the prompt.


LEARN

📹 Videos

✨ Quick links


USE CASES

🎬 Video Editors & Filmmakers
Generate voiceover, narration, or trailer reads for video projects without booking a voice actor.

🎙️ Podcast & Audio Content Creators
Produce two-speaker dialogue in a single run for interview-format intros, scripted segments, or explainer audio.

🎮 Game Developers
Create placeholder or production dialogue for NPCs, cutscenes, and tutorials with per-line emotion control.

📢 Marketing & Ad Teams
Turn ad scripts into voiced reads across different tones and styles, then compare variations before committing to a final cut.

📚 E-Learning & Accessibility
Voice course content, onboarding guides, or documentation for learners who prefer audio over text.


WHAT WORKS BEST / WHAT TO AVOID

✅ Works great

  • Clear, well-punctuated scripts

  • Short style phrases ("warm and playful," "whispered urgently")

  • Two-speaker dialogue with distinct names

  • Inline vocal sounds like <laugh> and <short pause>

⚠️ May produce softer results

  • Instructions written inside the prompt (the model reads them aloud)

  • More than two speakers (not supported)

  • Sound effects, music, or non-vocal sounds

  • Speaker names that differ only by capitalization ("Sam" vs "sam")


FAQ

What is Gemini 3.8 Flash-Lite TTS?
Gemini 3.8 Flash-Lite TTS is Google's fast, cost-efficient text-to-speech model from the Gemini 3.8 family. It supports 30 prebuilt voices, over 100 languages, style instructions, and two-speaker dialogue. It is the lighter counterpart to Flash TTS, optimized for high-volume production like dubbing and voice agents.

What is the difference between Gemini 3.8 Flash TTS and Flash-Lite TTS?
Both share the same settings and voice library, so you can swap between them without changing anything. Flash covers more languages and ranks higher on voice quality. Lite is lighter and faster, built for volume. Use Flash for character work or long-form narration where expressive depth matters. Use Lite for production at scale.

Can Gemini 3.8 Flash-Lite TTS do two-speaker dialogue?
Yes. Fill in both speaker ID fields, pick a voice for each, and write your script as Name: what they say. The whole conversation generates in a single run. No separate runs, no audio stitching. The two names must be different, and the model supports exactly two speakers per run.

How do I control emotion and pacing in the voice?
Use style instructions for the overall tone of the read, like "deep gravitas, slow cinematic pacing." For a single line, put the style in square brackets before the colon: Maya [genuinely surprised]: Wait, it works? Add timing with <short pause> and <long pause> tags, and vocal sounds like <laugh> or <sigh> inline in the script.

Does Gemini 3.8 Flash-Lite TTS support custom voice cloning?
No. The model offers 30 fixed prebuilt voices. There is no option to upload or clone a custom voice. You control the sound through voice selection and style instructions.

Can I use the audio output commercially?
Gemini 3.8 Flash-Lite TTS is a proprietary Google model. Commercial use is governed by Google's terms. Review Google's current usage policy for generated audio before using outputs in shipped products.

How to run Gemini 3.8 Flash-Lite TTS online?
You can run Gemini 3.8 Flash-Lite TTS online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, type your script, and hit run. Free to try.


WHY FLOYO?

Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.

A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.

For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.


Ready to try it?
Type your first script, pick a voice, and run it. The style is already set.

→ Launch Workflow, Free

Questions? Watch the free course or check the FAQ above.

Read more

N