Gemini 3.8 Flash TTS: Design Voices From Text Prompts

Gemini 3.8 Flash TTS: Design Voices From Text Prompts

  • News
  • Rocks on Galaxy
  • Apps
  • 25 Sep, 2026
  • 0

Google turned TTS from a preset picker into a voice studio. On September 23, 2026, the company introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS in a Google blog post by Leland Rechis and Alan Cowen. Flash is pitched for character design and line-by-line direction; Flash-Lite for high-volume dubbing and voice agents.

What you can actually do today

Flash TTS can create voices from natural-language prompts (role, accent, characteristics) across more than 100 languages and dialects, and it exposes a library of 2,000+ production-ready voices including regional varieties such as Mexican Spanish, Quebec French, and Scots English. Voice replication rebuilds a profile from a 30-second sample of a voice you have rights to use, with verbal consent matching the reference speaker before creation. Every Gemini Audio clip is watermarked with SynthID; C2PA credentials are also cited. Native two-speaker scene staging lets a single script drive a conversation with separated voices; long-form generation is described as holding timbre across hours with minimal drift.

  • Flash TTS: Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, Google Vids
  • Flash-Lite TTS: Gemini API and AI Studio now; Enterprise “coming soon”; Google Vids for everyone
  • Replication geo limits: not available in Illinois, Texas, EEA, UK, Switzerland, and India via AI Studio
  • Remixing: fine-tune library voices (timbre/pitch/pace/accent) listed as coming soon

Benchmarks—vendor-cited

Google says Flash TTS took #1 overall on Hume AI’s Voice Design Benchmark (71.4) and accent modeling (60.8), with Flash and Flash-Lite #1 and #2 on Hume’s Overall Quality Index. Those are scores Google published, not an independent Rocks lab. Voice Arena blind prefs are cited for Japanese, Brazilian Portuguese, Vietnamese, MSA Arabic, Mexican Spanish, and Hindi.

Studio workflow and partners

Google’s AI Studio audio playground is framed as a voice-design workspace: prompt new vocal identities or replicate your own voice (where allowed), then drop into a dual-speaker screenplay editor for line-by-line direction. Scripted vocal bursts and backchanneling (e.g., active-listening interjections) are listed for conversational texture. Models complement the broader Gemini Audio family after 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking. Developer platforms named as enabling speech experiences via the Gemini API include Agora, LiveKit, Pipecat, and Vercel. Integration partners cited: Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang for dubbing, accent localization, and conversational agents. Flash-Lite is optimized for high-volume dubbing and expressive voice agents with fine control over tone and pacing—still a Google product description, not a Rocks latency test.

Rocks take

Dated product launch with explicit consent and watermark language—and explicit places where cloning is blocked. Do not invent dollar-per-million API prices from this blog post; it does not list them. Source: Google.

FAQ

Can anyone clone any voice from 30 seconds? Google requires a matching verbal consent recording from the voice owner; AI Studio replication is blocked in several regions listed above.

Are Flash and Flash-Lite the same model? No. Flash is the creative/direction SKU; Flash-Lite is the volume/cost SKU.