Aggregate Edge-TTS, Kokoro, CosyVoice, and OpenAI into a single, blazing-fast self-hosted speech synthesis API gateway.
High concurrency, sub-200ms streaming latency, and universal client compatibility.
Direct drop-in replacement for standard OpenAI /v1/audio/speech, /v1/models, and /v1/audio/voices endpoints.
Pipeline sentence chunking and HTTP chunked transfer encoding deliver instantaneous audio playback without waiting for full generation.
Unified interface for Microsoft Edge-TTS, Kokoro-82M ONNX, CosyVoice, and OpenAI upstream relays.
Automatic MD5 disk caching prevents duplicate synthesis requests and yields sub-10ms response times for recurring phrases.
Deploy on your own infrastructure with single-port Docker or uv. Zero audio data leakage to external clouds.
Full support for custom pitch, speed, volume, and multi-format output (MP3, WAV, AAC, OPUS, FLAC).
Choose the optimal engine according to your speed, hardware, and voice quality requirements.
| Engine | Type | Voices / Languages | Latency | Hardware | Best For |
|---|---|---|---|---|---|
| Edge-TTSRecommended | Cloud / Free | 400+ Voices / 40+ Langs | 150 - 300ms | Low CPU / Memory | Everyday TTS, multi-lingual reading, audiobook production |
| Kokoro-82MUltra Fast | Local ONNX | English / Japanese / Chinese | < 150ms | CPU or CUDA GPU | Real-time conversational voice agents, gaming NPC voiceover |
| CosyVoice | Local Neural | Zero-shot Voice Clone | 250 - 450ms | NVIDIA GPU (CUDA) | High-fidelity personalized voice cloning, emotional speech |
| OpenAI Relay | Cloud Upstream | alloy, echo, fable, onyx, nova, shimmer | 300 - 600ms | API Key only | OpenAI official fallback & multi-model proxying |
From text request to audio stream in four streamlined processing steps.
Validates API key (Bearer/X-API-Key), parses parameters (model, voice, input, speed, format).
Computes query MD5 hash to return cached audio instantly, or segments text into natural phonetic chunks.
Dispatches chunks to the designated engine (EdgeTTS, Kokoro, CosyVoice) with parallel synthesis.
Encodes output format (mp3/wav/aac/opus) and streams chunks directly to client while caching result.
Seamlessly switch from OpenAI TTS to OneTTS by changing only base_url.
from openai import OpenAI
# Initialize client pointing to your local OneTTS gateway
client = OpenAI(
base_url="http://localhost:8030/v1",
api_key="sk-onetts-v1-k_8LxQm9Z2sTwE7rY1u"
)
# Synthesize speech stream using Edge-TTS or Kokoro
response = client.audio.speech.create(
model="edgetts",
voice="zh-CN-YunxiNeural",
input="你好,这是通过 OneTTS 高性能语音合成网关生成的高保真音频流!",
response_format="mp3",
speed=1.0
)
# Stream or save audio directly
response.stream_to_file("output.mp3")