Speech Synthesis API (POST /v1/audio/speech)
Generates audio from text matching the OpenAI /v1/audio/speech specification.
📥 Request Body
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
model | string | Optional | edgetts | Model or Provider ID (tts-1, edgetts, kokoro) |
input | string | Yes | - | Text to synthesize (up to 8,192 characters) |
voice | string | Optional | alloy | Voice ID or alias (alloy, zh-CN-XiaoxiaoNeural) |
response_format | string | Optional | mp3 | Format: mp3, opus, aac, flac, wav, pcm |
speed | float | Optional | 1.0 | Speed rate multiplier (0.25 to 4.0) |
stream | boolean | Optional | false | Enable HTTP Chunked audio streaming |
💡 Example
bash
curl http://localhost:8030/v1/audio/speech \
-H "Authorization: Bearer sk-onetts-v1-k9L3mX8QZ2sT4vA7wE1rY6u" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"input": "OneTTS delivers ultra-fast speech synthesis.",
"voice": "alloy",
"response_format": "mp3"
}' \
--output speech.mp3