Google Gemini 3.8 TTS: 30-second sample can replicate a voice
2026-09-24 08:38:37
According to CoinMeta, Google has made available two text-to-speech models, Gemini 3.8, TTS, and Flash - Lite TTS, in Gemini API and Google AI Studio. The Flash model focuses on sound quality, character performance, and stability with long texts, while the Lite model emphasizes high throughput, low latency, and low cost. Users can use a reference audio of about 30 seconds to replicate their own voice or an authorized voice, and control emotions, rhythm, and accent sentence by sentence. They can also insert sounds such as laughter, sighs, and coughs. Flash supports 130 languages, and Lite supports 101 languages. As of December 31, 2026, the cost of audio output with Flash is $9 per million tokens, and with Lite it is $6. Starting from January 1, 2027, the standard prices for both models will rise to $18 and $12 respectively.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.