← Back to all posts
News

Google's Gemini 3.8 TTS Adds 30-Second Voice Cloning at $9 per Million Audio Tokens

September 24, 2026 · 07:18 UTC · News
Google's Gemini 3.8 TTS Adds 30-Second Voice Cloning at $9 per Million Audio Tokens

TL;DR

On September 23 Google shipped Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS in the Gemini API and Google AI Studio. The feature that matters is voice replication: a 10-to-30-second reference clip plus a same-length recording of the same person reading a consent sentence becomes a reusable voice ID. Audio output costs $9 per million tokens on Flash TTS and $6 on Flash-Lite through December 31, 2026, then doubles on January 1, 2027; the 3.1 Flash TTS preview it replaces charges $20. Replication is unavailable in Illinois, Texas, the EEA, the UK, Switzerland and India. And the benchmark that crowns the model was built by Hume AI, whose founder now co-signs Google's launch post.


What shipped

Two models, one API surface. Flash TTS is the "creative direction" model: line-by-line performance cues, scripted laughs and sighs, two-speaker scenes from a single script, 130 languages with automatic detection. Flash-Lite TTS is the volume model for dubbing and voice agents, with 101 languages. Both do voice design from a text description, both do replication from a recording, and both are live today in the API and AI Studio, with Gemini Enterprise access "coming soon." Flash TTS is also wired into Gemini Notebook, Flash-Lite into Google Vids.

The launch post promises you can "scale up from 30 original voices to an infinite library." The speech generation docs put infinity at 200 stored voices per project, plus as many seven-day stateless keys as you care to manage yourself. Google also ships 2,000-plus prebuilt voices with regional varieties like Mexican Spanish, Quebec French and Scots English, so most teams will never hit the cap. But if you were planning a per-customer voice for a 10,000-seat product, read the limits page before the pitch deck.

The price, and the clock on it

The pricing page is unusually explicit about the promotional window. Standard Flash TTS is $0.50 per million input text tokens and $9.00 per million audio output tokens "through December 31, 2026," then $1.00 and $18.00 "starting January 1, 2027." Flash-Lite TTS is $0.50 and $6.00 now, $1.00 and $12.00 next year. Batch halves every number. The outgoing 3.1 Flash TTS preview sits at $1.00 in and $20.00 out with no promotion.

audio output, $ per 1M tokens, standard tier (lower is better) 3.1 Flash TTS preview$20 3.8 Flash TTS, 2027$18 3.8 Flash-Lite TTS, 2027$12 3.8 Flash TTS, 2026$9 3.8 Flash-Lite TTS, 2026$6
Through December 31 the new models cost $9 and $6 per million audio tokens; on January 1 both double, still under the outgoing preview's $20.

Two things to note before you model costs. Google prices audio by token, not by character or by minute, and the pricing page does not publish a tokens-per-second figure for these models, so any per-hour number you see this week is somebody's arithmetic, not Google's. And the doubling is on the calendar, not tied to usage, so a voice agent you launch in November costs twice as much in January without you touching a line.

How replication works

The voice replication docs describe a two-file handshake. You upload a 10-to-30-second reference clip of "clean, natural speech" and a consent clip of the same length in which the same adult speaker says, verbatim: "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model," or its equivalent in one of 30 supported languages. The API checks that the two recordings come from the same speaker before it builds anything. Google recommends the same microphone and the same room for both clips, because the check is acoustic, not legal.

Think of it as a notary who compares the signature on the form with the one on your ID: it proves the person consenting is the person on the tape, and nothing more. It does not know whether the voice on both clips is actually yours, and it does not know whether the person on the tape understood what they were agreeing to.

10-30 s clipreference speech consent clipsame length speaker matchsame adult voice voice ID1 yr / 7 d stored voice_ id: 1-year TTL, 200 per project stateless voicekey_: encrypted, 7-day TTL, you hold it
Two clips in, one voice handle out. Every generated clip carries a SynthID watermark; replicated voices also get a C2PA credential.

You choose where the voice lives. With store set to true you get a persistent voice ID that Google keeps for a year, capped at 200 per project. With store set to false you get an encrypted key that expires in seven days and that you manage yourself, which is the option for anyone whose lawyers do not want customer voiceprints sitting in a Google project. Output comes back as WAV, raw 16-bit PCM, mu-law or A-law at 24, 16 or 8 kHz, so the telephony crowd is covered. Every clip is watermarked with SynthID, and replicated voices additionally carry C2PA content credentials.

Six places where the clone button is missing

One sentence in the launch post does a lot of work: "Voice replication through AI Studio is not available in Illinois, Texas, EEA, UK, Switzerland, and India." Google does not say why. But the list reads like a map of biometric-consent law. Illinois' BIPA and Texas' CUBI both name voiceprints explicitly and BIPA carries a private right of action; GDPR treats voice data used to identify a person as a special category, and the UK and Switzerland run near-copies; India's DPDP Act arrived in 2023. The voice-replication docs and the pricing page I read do not repeat the region list, and the post's wording is specific to AI Studio, so whether the API path is gated the same way is a question for your account rep, not for this article.

The practical consequence for a small studio is real. If your narrator lives in Austin or Chicago, or your dubbing client is in Berlin, the cheapest cloning path on the market today is one you cannot legally use through the front door. Designed voices from a text prompt are not restricted, and for most product work that is the feature you actually wanted anyway.

The scoreboard, and who built it

The launch post claims "the #1 overall spot on Hume AI's Voice Design Benchmark (71.4)" and "#1 and #2 spots" on Hume's Overall Quality Index. The DeepMind speech page publishes the full tables. On voice design, Flash TTS scores 71.4 overall to ElevenLabs v3's 70.8 and Inworld's 69.8, and it runs away on accents at 60.8 against 45.4 and 35.8. On the "voice qualities" column it loses: 74.6 to ElevenLabs' 76.6 and Inworld's 76.3. On Hume's quality index, Flash TTS posts 0.920 and Flash-Lite 0.914, with Cartesia Sonic 3.6 at 0.840, the old 3.1 Flash TTS at 0.783, OpenAI's gpt-4o-mini-tts at 0.740 and ElevenLabs v3 at 0.706, though ElevenLabs takes a perfect 5.00 on human-like variation.

Gemini 3.8 Flash TTS ElevenLabs v3 Inworld 71.4 70.8 69.8 overall English 60.8 45.4 35.8 accents 74.6 76.6 76.3 voice qualities Hume AI Voice Design Benchmark, as published on Google's DeepMind speech page
Google leads overall and by a mile on accents, but ElevenLabs and Inworld both edge it on voice qualities.

Now the byline. The post is signed by Leland Rechis, Group Product Manager, and by "Alan Cowen, Director, Research Science, on Behalf of the Gemini Audio Team." Cowen founded Hume AI. On January 22 TechCrunch, citing WIRED, reported that Google DeepMind hired Cowen and roughly seven Hume engineers as part of a non-exclusive licensing deal; Hume stayed independent under new CEO Andrew Ettinger and kept selling to other labs. None of that makes the numbers wrong, and Hume's Real World VoiceEQ bench is public. It does mean the launch post's headline benchmark was built by a company the co-author founded and Google now licenses from, and the post does not mention it. When the grader's founder signs the report card, you read the third-party column first.

That column exists. Voice Arena runs blind, human, pairwise comparisons scored with Bradley-Terry Elo. On US English, Flash-Lite TTS is first at 1087, Cartesia's Sonic-3.6 second at 1068, Flash TTS third at 1061, Inworld's Realtime TTS 2 fourth at 1058, and the old 3.1 Flash TTS fifth at 1057. ElevenLabs' Eleven v3 Conversational is eighth at 1031. Google's own page adds that the 3.8 models take the top slot in Japanese, Brazilian Portuguese, Vietnamese, Arabic, Hindi and Mexican Spanish.

Voice Arena, US English, blind pairwise Elo (axis starts at 1000) Gemini 3.8 Flash-Lite TTS1087 Cartesia Sonic-3.61068 Gemini 3.8 Flash TTS1061 Inworld Realtime TTS 21058 Gemini 3.1 Flash TTS1057 Eleven v3 Conversational1031
The cheaper Google model tops the human-judged US English arena; the ElevenLabs entry sits 56 Elo back in eighth.

Read the Elo gaps with the error bars in mind: Voice Arena lists plus or minus 15 for Flash-Lite and 12 for Flash TTS, so first, second and third are within noise of each other. The old 3.1 model at fifth is the more telling number: one model generation moved the arena score by 30 points and cut the 2026 price by more than half.

What the first day of testers found

The Hacker News thread hit 296 points and 132 comments in its first day. The recurring quality complaint is overacting: the models "exaggerate all the tone and trailing" expressive sounds, in one commenter's words, which is the flip side of topping an expressiveness benchmark. Several people in the restricted regions found the clone tab simply absent. And the local-model crowd showed up with the usual list: Qwen3 TTS, Kokoro, Fish Audio, all of which run on a workstation and none of which double in price on New Year's Day. If you run your own stack, our earlier piece on Nari Labs' Qwen3-TTS serving stack is the counterpoint to this whole post.

Caveats before you switch

  • Requests are capped at 8,192 input tokens, and multi-speaker mode takes at most two speakers using prebuilt voices, so long dialogue scripts get chunked and stitched on your side.
  • The consent check verifies speaker identity between two clips. It is not proof of ownership and it is not a rights clearance; your contract with the voice actor still has to say the words.
  • The region list is quoted for AI Studio. Confirm API behavior for your jurisdiction before you build a product on replication, and assume your customers' locations matter, not just yours.
  • Pricing is per audio token with no published tokens-per-second figure, and both models double in price on January 1, 2027. Budget on the 2027 numbers.
  • Voice remixing (timbre, pitch, pace, accent) is listed as coming soon, not shipped.

Key Takeaways

  • Gemini 3.8 Flash TTS and Flash-Lite TTS launched September 23 in the Gemini API and AI Studio, with voice design, two-speaker scenes, and 130 and 101 languages respectively.
  • Voice replication takes a 10-to-30-second reference clip plus a same-length consent clip of the same adult speaker reading a fixed sentence; the API matches the two speakers before creating a voice.
  • Audio output is $9 per million tokens on Flash TTS and $6 on Flash-Lite through December 31, 2026, doubling to $18 and $12 on January 1, 2027; the 3.1 preview charges $20.
  • Replication through AI Studio is off in Illinois, Texas, the EEA, the UK, Switzerland and India, a list that tracks biometric-consent law; designed voices are unrestricted.
  • Google's headline benchmark wins come from Hume AI, whose founder Alan Cowen joined DeepMind in a January licensing deal and co-signs the launch post; the independent Voice Arena still puts Flash-Lite first on US English at 1087 Elo.
  • Stored voices cap at 200 per project with a one-year TTL; stateless seven-day keys let you keep voiceprints out of Google's storage.

Sources: Google: Gemini 3.8 Flash TTS and Flash-Lite TTS, Gemini API pricing, Gemini API voice replication docs, Gemini API speech generation docs, Gemini 3.8 Flash TTS model page, DeepMind: Gemini Audio speech generation benchmarks, Voice Arena US English leaderboard, TechCrunch: Google hires Hume AI team (Jan 22, 2026), Hacker News discussion

AIGoogleGeminitext-to-speechvoice cloningSynthIDElevenLabsbiometrics
CONSOLE
$