Skip to main content
This page extends the Gemini native format section of the Audio page and covers the new-generation TTS models.

Authentication

Gemini native format (TTS)

{model} is gemini-3.8-flash-tts or gemini-3.1-flash-tts-preview.

gemini-3.8-flash-tts

gemini-3.1-flash-tts-preview

Response

Handling the response

The audio is returned in candidates[0].content.parts[0].inlineData:
Adding a WAV header for 3.1 (Python; parse the sample rate from the response mimeType, donโ€™t hard-code it):

Parameters

Inline tags (shared by both generations, placed in the text to control momentary events): <laugh>, <sigh>, <cough>, <breath>, <short pause>. Multi-turn: 3.8 supports stable long multi-turn conversations; 3.1 is single-turn only โ€” contents may contain just one user turn, and multiple turns return Multiturn chat is not enabled for this model. Use gemini-3.8-flash-tts if you need multi-turn. Multi-speaker (3.8): each part carries its own speech_metadata.speaker, and speakers must match those configured in speechConfig:

Voice library lookup

Lists available voices (prebuilt / extended voice library). ?model= selects the channel the request is routed to. Lookups are not billed.

Request parameters

Request example

Response

The voice library covers 130+ languages, including regional variants such as ar-001 (Modern Standard Arabic) and ar-EG (Egyptian Arabic).

Model comparison

Both models support voice library lookup (/v1beta/voices) and prebuilt voices (voiceName). Input / output token limits: 8,192 / 16,384 for both (16,384 output tokens is roughly 11 minutes of audio).

Billing

Billed by actual usage, based on the response usageMetadata: Voice library lookups (/v1beta/voices) are not billed.

Common errors