Two voice models, four places to use them

Gemini Audio metacard graphic from Google's Gemini 3.8 text-to-speech announcement Caption: Metacard published with Google's 23 Sep 2026 Gemini TTS post · Source: Google Keyword blog · link

On 23 September 2026 Google published Gemini 3.8 text-to-speech says hello. The post ships two speech models: Gemini 3.8 Flash TTS (creative direction and character design) and Gemini 3.8 Flash-Lite TTS (high-volume, cost-efficient scale). Model codes in the Gemini API docs are gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.

Rollout detail in the post: both models are available today for developers on the Gemini API and Google AI Studio. Enterprise API access for both is coming soon via Gemini Enterprise. Flash's consumer path is Gemini Notebook; Flash-Lite's is Google Vids. The models sit beside the existing Gemini Audio line (3.5 Live Translate / Transcribe, 3.8 Live, 3.8 Live Extended Thinking).

Who gets it today and who waits

Google's product launch sets availability and region limits. It does not create a new statutory duty by itself.

  • Developers and voice-agent builders on Gemini API / AI Studio decide whether to migrate prompts from gemini-3.1-flash-tts-preview.
  • Enterprises waiting on Gemini Enterprise API access for either TTS model stay on the "coming soon" clock Google stated.
  • Media, dubbing, and audiobook teams decide whether generative voice design and 30-second voice replication fit rights workflows.
  • Operators in Illinois, Texas, the EEA, the UK, Switzerland, and India cannot use voice replication through AI Studio under the geographic limits Google listed.
  • Safety / trust teams own SynthID watermarking, C2PA credentials, and consent-recording checks for replicated voices.

Voice design, 2,000 voices and 30-second cloning

  • Natural-language generative voice design on Flash TTS: role, accent, and character traits. The Keyword post says more than 100 languages and dialects; the model card lists 130 languages for Flash TTS and 101 for Flash-Lite TTS.
  • Library scale claimed at 2,000+ voices, including regional varieties such as Mexican Spanish, Quebec French, and Scots English.
  • Voice replication from about 30 seconds of consented reference audio, with verbal consent verification, SynthID watermarking, and C2PA credentials.
  • Line-by-line performance direction, long-form generation with claimed low speaker drift, native two-speaker scene staging, and scripted vocal bursts / backchanneling tokens.
  • Google cites Hume AI Voice Design Benchmark leadership for Flash TTS (71.4 overall; 60.8 accent modeling) and #1 / #2 spots on Hume's Overall Quality Index for Flash and Flash-Lite. Blind Voice Arena preferences are claimed for several languages including Japanese, Brazilian Portuguese, Vietnamese, MSA Arabic, Mexican Spanish, and Hindi. The post compares gains against prior Gemini 3.1 Flash TTS.
  • Partner names on the post include developer platforms Agora, LiveKit, Pipecat, and Vercel, plus Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang.
  • Paid API pricing (standard tier, per 1M tokens, through 31 Dec 2026) is published on the Gemini API pricing page: Flash TTS $0.50 text input / $9.00 audio output; Flash-Lite TTS $0.50 text input / $6.00 audio output. From 1 Jan 2027 those rates rise to $1.00 / $18.00 (Flash) and $1.00 / $12.00 (Flash-Lite). Free-tier input and output are listed as free of charge. Google points readers to a model card for safety detail.

No rate limits and no enterprise date

  • Rate limits and latency SLOs are still not spelled out in the Keyword post. Check live quotas before forecasting capacity.
  • "Most expressive" and Hume / Voice Arena ranks are Google-reported or third-party leaderboard snapshots. They are not your A/B test on your scripts.
  • Voice remixing (example prompt: "add subtle Southern US accent") is labeled coming soon. It is not listed as shipping in this post.
  • Gemini Enterprise API access for both TTS models remains coming soon. Developer API and AI Studio access is what ships on day one.
  • Voice replication remains blocked in the listed jurisdictions even when the base TTS models roll out.
  • Consent verification reduces some misuse classes. It does not settle deepfake liability or local biometric / voice laws.

Test in AI Studio before you migrate

  1. Smoke-test Flash TTS in AI Studio on one dual-speaker script and one long-form chapter before migrating production traffic.
  2. If you need voice cloning, confirm your users are outside the blocked regions and that you can store the required verbal consent recording.
  3. Keep SynthID / C2PA detection in the release checklist for any customer-facing audio.
  4. Budget from the pricing page. Pin gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts and the Dec 2026 promotional rates above. The Keyword post does not list unit prices.
  5. Do not conflate this ship with Gemini 3.8 Live Extended Thinking. That earlier Live voice product is a different surface from these TTS generators.
  6. Pull Hume numbers only as vendor-cited context. Run your own preference test on your languages and agent prompts.