> ## Documentation Index
> Fetch the complete documentation index at: https://docs.gravitex.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Gemini TTS

> gemini-3.8-flash-tts / gemini-3.1-flash-tts-preview 语音合成与音色库查询

本页为 [音频（Audio）](/cn/api-reference/endpoint/audio) 页 Gemini 原生格式的扩展，覆盖新一代 TTS 模型。

## 鉴权

```
Authorization: Bearer sk-xxxxxxxxxx
```

## Gemini 原生格式(TTS)

```
POST https://api.gravitex.ai/v1beta/models/{model}:generateContent
```

`{model}` 为 `gemini-3.8-flash-tts` 或 `gemini-3.1-flash-tts-preview`。

### gemini-3.8-flash-tts

```bash theme={null}
curl "https://api.gravitex.ai/v1beta/models/gemini-3.8-flash-tts:generateContent" \
  -H "Authorization: Bearer sk-xxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [
      {
        "role": "user",
        "parts": [
          {
            "text": "欢迎收听今天的科技早报。<short pause>第一条,研究人员宣布在室温超导材料上取得突破。",
            "speech_metadata": { "speaker": "Host", "style": "新闻播报,清晰而热情" }
          }
        ]
      }
    ],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "voiceConfig": {
          "prebuiltVoiceConfig": { "voiceName": "Kore" }
        }
      }
    }
  }'
```

### gemini-3.1-flash-tts-preview

```bash theme={null}
curl "https://api.gravitex.ai/v1beta/models/gemini-3.1-flash-tts-preview:generateContent" \
  -H "Authorization: Bearer sk-xxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [
      {
        "role": "user",
        "parts": [
          { "text": "用欢快的语气说:祝你今天过得愉快!" }
        ]
      }
    ],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "voiceConfig": {
          "prebuiltVoiceConfig": { "voiceName": "Kore" }
        }
      }
    }
  }'
```

### 响应

```json theme={null}
{
  "candidates": [
    {
      "content": {
        "role": "model",
        "parts": [
          { "inlineData": { "mimeType": "audio/wav", "data": "UklGRi4..." } }
        ]
      },
      "finishReason": "STOP"
    }
  ],
  "usageMetadata": {
    "promptTokenCount": 5,
    "candidatesTokenCount": 1680,
    "totalTokenCount": 1685,
    "candidatesTokensDetails": [ { "modality": "AUDIO", "tokenCount": 1680 } ]
  },
  "modelVersion": "gemini-3.8-flash-tts",
  "responseId": "xxxx",
  "createTime": "2026-10-10T10:00:00Z"
}
```

| 字段 | 说明 |
| - | - |
| `candidates[].content.parts[].inlineData.mimeType` | 音频格式,见下方处理表 |
| `candidates[].content.parts[].inlineData.data` | 音频 base64 |
| `candidates[].finishReason` | `STOP`(正常结束)或其他截断原因 |
| `usageMetadata.promptTokenCount` | 文本输入 token(计费依据) |
| `usageMetadata.candidatesTokenCount` | 音频输出 token(计费依据) |
| `usageMetadata.totalTokenCount` | 合计 |
| `usageMetadata.candidatesTokensDetails[].modality` | 模态明细,音频输出为 `AUDIO` |
| `modelVersion` / `responseId` / `createTime` | 实际模型版本 / 请求标识 / 时间 |

### 响应处理

音频在 `candidates[0].content.parts[0].inlineData` 中返回:

```bash theme={null}
echo "<base64_data>" | base64 --decode > out.wav
```

| 模型 | mimeType | 处理方式 |
| - | - | - |
| gemini-3.8-flash-tts | `audio/wav` | 解码后直接是完整 WAV,可播放 |
| gemini-3.1-flash-tts-preview | `audio/l16;codec=pcm;rate=24000` | 裸 PCM,需补 WAV 头;实际参数以响应的 `mimeType` 为准 |

3.1 补 WAV 头(Python,采样率从响应的 `mimeType` 解析,勿硬编码):

```python theme={null}
import base64, json, struct

resp = json.load(open("response.json"))
inline = resp["candidates"][0]["content"]["parts"][0]["inlineData"]
pcm = base64.b64decode(inline["data"])

# mimeType 形如 audio/l16;codec=pcm;rate=24000,从 rate= 解析
rate = 24000
for part in inline["mimeType"].split(";"):
    if part.strip().startswith("rate="):
        rate = int(part.strip()[5:])

header = (b"RIFF" + struct.pack("<I", 36 + len(pcm)) + b"WAVE"
          + b"fmt " + struct.pack("<IHHIIHH", 16, 1, 1, rate, rate * 2, 2, 16)
          + b"data" + struct.pack("<I", len(pcm)))
open("speech.wav", "wb").write(header + pcm)
```

### 参数说明

| 参数 | 说明 |
| - | - |
| `contents[].parts[].text` | 要朗读的文本;3.1 风格指令写在此处(如"用欢快的语气说:") |
| `contents[].parts[].speech_metadata.speaker` | 说话人标签(仅 3.8;多说话人每轮必填) |
| `contents[].parts[].speech_metadata.style` | 持续朗读风格(仅 3.8;耳语、语速等) |
| `generationConfig.responseModalities` | 固定 `["AUDIO"]` |
| `generationConfig.speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName` | 音色 ID,从音色库查询获取 |
| `generationConfig.speechConfig.multiSpeakerVoiceConfig` | 多说话人配置 |

**内嵌标记**(两代通用,写在正文中控制瞬时事件):`<laugh>`、`<sigh>`、`<cough>`、`<breath>`、`<short pause>`。

**多轮:** 3.8 支持长对话多轮稳定;3.1 仅单轮——`contents` 只能有一个 user turn,多轮返回 `Multiturn chat is not enabled for this model`,需要多轮请使用 `gemini-3.8-flash-tts`。

**多说话人(3.8):** 每轮 part 各自带 `speech_metadata.speaker`,说话人需与 `speechConfig` 配置的说话人匹配:

```json theme={null}
{
  "contents": [
    {
      "role": "user",
      "parts": [
        { "text": "今晚的主题是人工智能安全。", "speech_metadata": { "speaker": "Alice", "style": "沉稳的主持人" } },
        { "text": "我先说结论:<short pause>最大的风险在部署环节。", "speech_metadata": { "speaker": "Bob", "style": "轻快的学者" } }
      ]
    }
  ],
  "generationConfig": {
    "responseModalities": ["AUDIO"],
    "speechConfig": {
      "voiceConfig": { "prebuiltVoiceConfig": { "voiceName": "Puck" } }
    }
  }
}
```

## 音色库查询

```
GET https://api.gravitex.ai/v1beta/voices?model={model}
```

查询可用音色(预置 / 扩展音色库),按 `?model=` 指定模型路由渠道。查询不产生计费。

### 请求参数

| 参数 | 必填 | 说明 |
| - | - | - |
| `model` | 是 | 模型名,如 `gemini-3.8-flash-tts` |
| `language_code` | 否 | 按语言过滤,可重复,如 `zh-CN`、`en-US`、`ar-EG` |
| `gender` | 否 | `female` / `male` |
| `pitch` | 否 | `low` / `medium` / `high` |
| `context` | 否 | 按适用场景过滤,取值与响应 `context` 字段对应:`Content & Media`、`Conversational / Edu`、`Enterprise Agent`、`Growth & Marketing`、`Entertainment & Gaming`、`Wellness & Culture` |
| `type` | 否 | `prebuilt`(预置)等 |
| `search` | 否 | 自由文本搜索音色特征,如 `warm` |
| `page_size` | 否 | 分页大小,默认 50 |
| `page_token` | 否 | 翻页令牌,取上一页响应的 `next_page_token` |

### 请求示例

```bash theme={null}
curl -G "https://api.gravitex.ai/v1beta/voices?model=gemini-3.1-flash-tts-preview" \
  -H "Authorization: Bearer sk-xxxxxxxxxx" \
  --data-urlencode "language_code=zh-CN" \
  --data-urlencode "gender=female" \
  --data-urlencode "search=warm"
```

### 响应

```json theme={null}
{
  "voices": [
    {
      "id": "achernar",
      "type": "VOICE_TYPE_PREBUILT",
      "display_name": "Achernar",
      "language_code": "en-US",
      "region_code": "US",
      "accent": "General American",
      "persona": "Storyteller & Narrator",
      "context": "Content & Media",
      "gender": "female",
      "pitch": "PITCH_HIGH",
      "description": "Soft, calm, and soothing voice with a higher pitch. Recommended for quiet or personal storytelling."
    },
    {
      "id": "achird",
      "type": "VOICE_TYPE_PREBUILT",
      "display_name": "Achird",
      "language_code": "en-US",
      "region_code": "US",
      "accent": "General American",
      "persona": "Companion & Peer",
      "context": "Conversational / Edu",
      "gender": "male",
      "pitch": "PITCH_LOW",
      "description": "Friendly, approachable, and warm voice with a lower-middle pitch. Recommended for casual walkthroughs or vlogs."
    }
  ],
  "next_page_token": "achird|en-US"
}
```

| 字段 | 说明 |
| - | - |
| `id` | 音色 ID,填入 `voiceName` 使用 |
| `type` | 音色类型,如 `VOICE_TYPE_PREBUILT`(预置) |
| `display_name` | 显示名称 |
| `language_code` / `region_code` | 语言与地区,如 `en-US` / `US` |
| `accent` | 口音,如 `General American` |
| `persona` | 人设,如 `Storyteller & Narrator`、`High-Trust Advisor` |
| `context` | 适用场景 |
| `gender` | `female` / `male` |
| `pitch` | `PITCH_HIGH` / `PITCH_MEDIUM` / `PITCH_LOW` |
| `description` | 音色描述 |
| `next_page_token` | 下一页令牌,传入下次请求的 `page_token`;无更多结果时省略 |

音色库覆盖 130+ 种语言(含区域变体,如 `ar-001` 现代标准阿拉伯语、`ar-EG` 埃及阿拉伯语)。

## 模型对照

| 模型 | 风格控制 | 多轮 | 输出格式 | 上下文缓存 | 语言 |
| - | - | - | - | - | - |
| `gemini-3.8-flash-tts` | `speech_metadata` 结构化 | ✅ | WAV | ✅ | 130+ |
| `gemini-3.1-flash-tts-preview` | 正文内嵌指令 | ❌ | 裸 PCM | ❌ | 多语言 |

两个模型均支持音色库查询(`/v1beta/voices`)与预置音色(`voiceName`)。

输入 / 输出 token 上限:均为 8,192 / 16,384(16,384 输出 token 约合 11 分钟音频)。

## 计费

按实际用量结算,以响应的 `usageMetadata` 为准:

| 项 | 价格 |
| - | - |
| 输入(文本) | \$0.50 / 1M tokens(`promptTokenCount`) |
| 输出(音频) | \$9.00 / 1M tokens(`candidatesTokenCount`;约每 10 秒音频 250 token) |

音色库查询(`/v1beta/voices`)不计费。

## 常见错误

| 错误 | 原因与处理 |
| - | - |
| `Multiturn chat is not enabled for this model` | 3.1 仅支持单轮,把全部内容合并进一个 user turn;需要多轮换 `gemini-3.8-flash-tts` |
| `Invalid token` | API Key 无效或未携带 `Authorization: Bearer` 头 |
| `No available channel for model ...` | 模型名有误(如未替换 `{model}` 占位符),或当前分组无可用渠道 |
| 生成被截断 | 输出超 16,384 token(约 11 分钟),分段合成 |

## 相关链接

* [Google Speech Generation 文档](https://ai.google.dev/gemini-api/docs/speech-generation?hl=zh-cn)
* [gemini-3.8-flash-tts 模型页](https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts?hl=zh-cn)
* [gemini-3.1-flash-tts-preview 模型页](https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-tts-preview?hl=zh-cn)
* [GravitexAI 音频总览](https://docs.gravitex.ai/cn/api-reference/endpoint/audio)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.