> ## Documentation Index
> Fetch the complete documentation index at: https://docs.gravitex.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Gemini TTS

> Speech synthesis and voice library lookup for gemini-3.8-flash-tts / gemini-3.1-flash-tts-preview

This page extends the Gemini native format section of the [Audio](/en/api-reference/endpoint/audio) page and covers the new-generation TTS models.

## Authentication

```
Authorization: Bearer sk-xxxxxxxxxx
```

## Gemini native format (TTS)

```
POST https://api.gravitex.ai/v1beta/models/{model}:generateContent
```

`{model}` is `gemini-3.8-flash-tts` or `gemini-3.1-flash-tts-preview`.

### gemini-3.8-flash-tts

```bash theme={null}
curl "https://api.gravitex.ai/v1beta/models/gemini-3.8-flash-tts:generateContent" \
  -H "Authorization: Bearer sk-xxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [
      {
        "role": "user",
        "parts": [
          {
            "text": "Welcome to the daily tech briefing.<short pause>First, researchers announce a breakthrough in room-temperature superconducting materials.",
            "speech_metadata": { "speaker": "Host", "style": "News broadcast, clear and enthusiastic" }
          }
        ]
      }
    ],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "voiceConfig": {
          "prebuiltVoiceConfig": { "voiceName": "Kore" }
        }
      }
    }
  }'
```

### gemini-3.1-flash-tts-preview

```bash theme={null}
curl "https://api.gravitex.ai/v1beta/models/gemini-3.1-flash-tts-preview:generateContent" \
  -H "Authorization: Bearer sk-xxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [
      {
        "role": "user",
        "parts": [
          { "text": "Say cheerfully: Have a wonderful day!" }
        ]
      }
    ],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "voiceConfig": {
          "prebuiltVoiceConfig": { "voiceName": "Kore" }
        }
      }
    }
  }'
```

### Response

```json theme={null}
{
  "candidates": [
    {
      "content": {
        "role": "model",
        "parts": [
          { "inlineData": { "mimeType": "audio/wav", "data": "UklGRi4..." } }
        ]
      },
      "finishReason": "STOP"
    }
  ],
  "usageMetadata": {
    "promptTokenCount": 5,
    "candidatesTokenCount": 1680,
    "totalTokenCount": 1685,
    "candidatesTokensDetails": [ { "modality": "AUDIO", "tokenCount": 1680 } ]
  },
  "modelVersion": "gemini-3.8-flash-tts",
  "responseId": "xxxx",
  "createTime": "2026-10-10T10:00:00Z"
}
```

| Field | Description |
| - | - |
| `candidates[].content.parts[].inlineData.mimeType` | Audio format; see the handling table below |
| `candidates[].content.parts[].inlineData.data` | Base64-encoded audio |
| `candidates[].finishReason` | `STOP` (normal completion) or another truncation reason |
| `usageMetadata.promptTokenCount` | Text input tokens (billing basis) |
| `usageMetadata.candidatesTokenCount` | Audio output tokens (billing basis) |
| `usageMetadata.totalTokenCount` | Total |
| `usageMetadata.candidatesTokensDetails[].modality` | Per-modality breakdown; `AUDIO` for audio output |
| `modelVersion` / `responseId` / `createTime` | Actual model version / request ID / timestamp |

### Handling the response

The audio is returned in `candidates[0].content.parts[0].inlineData`:

```bash theme={null}
echo "<base64_data>" | base64 --decode > out.wav
```

| Model | mimeType | Handling |
| - | - | - |
| gemini-3.8-flash-tts | `audio/wav` | Decodes directly to a complete, playable WAV |
| gemini-3.1-flash-tts-preview | `audio/l16;codec=pcm;rate=24000` | Raw PCM; a WAV header must be added. Use the parameters in the response `mimeType` |

Adding a WAV header for 3.1 (Python; parse the sample rate from the response `mimeType`, don't hard-code it):

```python theme={null}
import base64, json, struct

resp = json.load(open("response.json"))
inline = resp["candidates"][0]["content"]["parts"][0]["inlineData"]
pcm = base64.b64decode(inline["data"])

# mimeType looks like audio/l16;codec=pcm;rate=24000; parse rate= from it
rate = 24000
for part in inline["mimeType"].split(";"):
    if part.strip().startswith("rate="):
        rate = int(part.strip()[5:])

header = (b"RIFF" + struct.pack("<I", 36 + len(pcm)) + b"WAVE"
          + b"fmt " + struct.pack("<IHHIIHH", 16, 1, 1, rate, rate * 2, 2, 16)
          + b"data" + struct.pack("<I", len(pcm)))
open("speech.wav", "wb").write(header + pcm)
```

### Parameters

| Parameter | Description |
| - | - |
| `contents[].parts[].text` | Text to read aloud; for 3.1, put style instructions here (e.g. "Say cheerfully:") |
| `contents[].parts[].speech_metadata.speaker` | Speaker label (3.8 only; required on every turn for multi-speaker) |
| `contents[].parts[].speech_metadata.style` | Persistent speaking style (3.8 only; whisper, pace, etc.) |
| `generationConfig.responseModalities` | Always `["AUDIO"]` |
| `generationConfig.speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName` | Voice ID, obtained from the voice library |
| `generationConfig.speechConfig.multiSpeakerVoiceConfig` | Multi-speaker configuration |

**Inline tags** (shared by both generations, placed in the text to control momentary events): `<laugh>`, `<sigh>`, `<cough>`, `<breath>`, `<short pause>`.

**Multi-turn:** 3.8 supports stable long multi-turn conversations; 3.1 is single-turn only — `contents` may contain just one user turn, and multiple turns return `Multiturn chat is not enabled for this model`. Use `gemini-3.8-flash-tts` if you need multi-turn.

**Multi-speaker (3.8):** each part carries its own `speech_metadata.speaker`, and speakers must match those configured in `speechConfig`:

```json theme={null}
{
  "contents": [
    {
      "role": "user",
      "parts": [
        { "text": "Tonight's topic is AI safety.", "speech_metadata": { "speaker": "Alice", "style": "Calm host" } },
        { "text": "Let me start with the conclusion:<short pause>the biggest risk is in deployment.", "speech_metadata": { "speaker": "Bob", "style": "Lively scholar" } }
      ]
    }
  ],
  "generationConfig": {
    "responseModalities": ["AUDIO"],
    "speechConfig": {
      "voiceConfig": { "prebuiltVoiceConfig": { "voiceName": "Puck" } }
    }
  }
}
```

## Voice library lookup

```
GET https://api.gravitex.ai/v1beta/voices?model={model}
```

Lists available voices (prebuilt / extended voice library). `?model=` selects the channel the request is routed to. Lookups are not billed.

### Request parameters

| Parameter | Required | Description |
| - | - | - |
| `model` | Yes | Model name, e.g. `gemini-3.8-flash-tts` |
| `language_code` | No | Filter by language; repeatable, e.g. `zh-CN`, `en-US`, `ar-EG` |
| `gender` | No | `female` / `male` |
| `pitch` | No | `low` / `medium` / `high` |
| `context` | No | Filter by use case, matching the response `context` field: `Content & Media`, `Conversational / Edu`, `Enterprise Agent`, `Growth & Marketing`, `Entertainment & Gaming`, `Wellness & Culture` |
| `type` | No | `prebuilt` etc. |
| `search` | No | Free-text search over voice traits, e.g. `warm` |
| `page_size` | No | Page size, default 50 |
| `page_token` | No | Pagination token; use `next_page_token` from the previous response |

### Request example

```bash theme={null}
curl -G "https://api.gravitex.ai/v1beta/voices?model=gemini-3.1-flash-tts-preview" \
  -H "Authorization: Bearer sk-xxxxxxxxxx" \
  --data-urlencode "language_code=zh-CN" \
  --data-urlencode "gender=female" \
  --data-urlencode "search=warm"
```

### Response

```json theme={null}
{
  "voices": [
    {
      "id": "achernar",
      "type": "VOICE_TYPE_PREBUILT",
      "display_name": "Achernar",
      "language_code": "en-US",
      "region_code": "US",
      "accent": "General American",
      "persona": "Storyteller & Narrator",
      "context": "Content & Media",
      "gender": "female",
      "pitch": "PITCH_HIGH",
      "description": "Soft, calm, and soothing voice with a higher pitch. Recommended for quiet or personal storytelling."
    },
    {
      "id": "achird",
      "type": "VOICE_TYPE_PREBUILT",
      "display_name": "Achird",
      "language_code": "en-US",
      "region_code": "US",
      "accent": "General American",
      "persona": "Companion & Peer",
      "context": "Conversational / Edu",
      "gender": "male",
      "pitch": "PITCH_LOW",
      "description": "Friendly, approachable, and warm voice with a lower-middle pitch. Recommended for casual walkthroughs or vlogs."
    }
  ],
  "next_page_token": "achird|en-US"
}
```

| Field | Description |
| - | - |
| `id` | Voice ID; pass it as `voiceName` |
| `type` | Voice type, e.g. `VOICE_TYPE_PREBUILT` (prebuilt) |
| `display_name` | Display name |
| `language_code` / `region_code` | Language and region, e.g. `en-US` / `US` |
| `accent` | Accent, e.g. `General American` |
| `persona` | Persona, e.g. `Storyteller & Narrator`, `High-Trust Advisor` |
| `context` | Suitable use case |
| `gender` | `female` / `male` |
| `pitch` | `PITCH_HIGH` / `PITCH_MEDIUM` / `PITCH_LOW` |
| `description` | Voice description |
| `next_page_token` | Next-page token; pass it as `page_token` in the next request. Omitted when there are no more results |

The voice library covers 130+ languages, including regional variants such as `ar-001` (Modern Standard Arabic) and `ar-EG` (Egyptian Arabic).

## Model comparison

| Model | Style control | Multi-turn | Output format | Context caching | Languages |
| - | - | - | - | - | - |
| `gemini-3.8-flash-tts` | Structured `speech_metadata` | ✅ | WAV | ✅ | 130+ |
| `gemini-3.1-flash-tts-preview` | Inline instructions in text | ❌ | Raw PCM | ❌ | Multilingual |

Both models support voice library lookup (`/v1beta/voices`) and prebuilt voices (`voiceName`).

Input / output token limits: 8,192 / 16,384 for both (16,384 output tokens is roughly 11 minutes of audio).

## Billing

Billed by actual usage, based on the response `usageMetadata`:

| Item | Price |
| - | - |
| Input (text) | \$0.50 / 1M tokens (`promptTokenCount`) |
| Output (audio) | \$9.00 / 1M tokens (`candidatesTokenCount`; about 250 tokens per 10 seconds of audio) |

Voice library lookups (`/v1beta/voices`) are not billed.

## Common errors

| Error | Cause and fix |
| - | - |
| `Multiturn chat is not enabled for this model` | 3.1 supports single-turn only; merge everything into one user turn, or switch to `gemini-3.8-flash-tts` for multi-turn |
| `Invalid token` | The API key is invalid or the `Authorization: Bearer` header is missing |
| `No available channel for model ...` | Wrong model name (e.g. the `{model}` placeholder wasn't replaced), or no channel is available for the current group |
| Output truncated | Output exceeded 16,384 tokens (about 11 minutes); synthesize in segments |

## Related links

* [Google Speech Generation docs](https://ai.google.dev/gemini-api/docs/speech-generation)
* [gemini-3.8-flash-tts model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts)
* [gemini-3.1-flash-tts-preview model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-tts-preview)
* [GravitexAI Audio overview](https://docs.gravitex.ai/en/api-reference/endpoint/audio)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.