Authentication
Gemini native format (TTS)
{model} is gemini-3.8-flash-tts or gemini-3.1-flash-tts-preview.
gemini-3.8-flash-tts
gemini-3.1-flash-tts-preview
Response
Handling the response
The audio is returned incandidates[0].content.parts[0].inlineData:
Adding a WAV header for 3.1 (Python; parse the sample rate from the response
mimeType, donโt hard-code it):
Parameters
Inline tags (shared by both generations, placed in the text to control momentary events):
<laugh>, <sigh>, <cough>, <breath>, <short pause>.
Multi-turn: 3.8 supports stable long multi-turn conversations; 3.1 is single-turn only โ contents may contain just one user turn, and multiple turns return Multiturn chat is not enabled for this model. Use gemini-3.8-flash-tts if you need multi-turn.
Multi-speaker (3.8): each part carries its own speech_metadata.speaker, and speakers must match those configured in speechConfig:
Voice library lookup
?model= selects the channel the request is routed to. Lookups are not billed.
Request parameters
Request example
Response
The voice library covers 130+ languages, including regional variants such as
ar-001 (Modern Standard Arabic) and ar-EG (Egyptian Arabic).
Model comparison
Both models support voice library lookup (
/v1beta/voices) and prebuilt voices (voiceName).
Input / output token limits: 8,192 / 16,384 for both (16,384 output tokens is roughly 11 minutes of audio).
Billing
Billed by actual usage, based on the responseusageMetadata:
Voice library lookups (
/v1beta/voices) are not billed.
