Skip to content
OneTTS Gateway · High Performance Speech Synthesis

Unified Speech SynthesisOpenAI-Compatible TTS Gateway

Aggregate Edge-TTS, Kokoro, CosyVoice, and OpenAI into a single, blazing-fast self-hosted speech synthesis API gateway.

# 1. Pull and run OneTTS Gateway Service
$ docker run -d -p 8030:8030 --name onetts codeandxv/onetts:latest
INFO: Registering TTS Engines: [edgetts, kokoro, openai]
INFO: Application startup complete. Uvicorn running on http://0.0.0.0:8030
OpenAI Speech API ready at /v1/audio/speech

Engineered for Real-Time Voice AI

High concurrency, sub-200ms streaming latency, and universal client compatibility.

🎙️Standard API

OpenAI Speech Compatible

Direct drop-in replacement for standard OpenAI /v1/audio/speech, /v1/models, and /v1/audio/voices endpoints.

⚡< 200ms

Sub-200ms Chunked Stream

Pipeline sentence chunking and HTTP chunked transfer encoding deliver instantaneous audio playback without waiting for full generation.

🧩Multi-Engine

Multi-Engine Aggregation

Unified interface for Microsoft Edge-TTS, Kokoro-82M ONNX, CosyVoice, and OpenAI upstream relays.

💾Cache

Smart MD5 Audio Caching

Automatic MD5 disk caching prevents duplicate synthesis requests and yields sub-10ms response times for recurring phrases.

🛡️Private

100% Self-Hosted & Private

Deploy on your own infrastructure with single-port Docker or uv. Zero audio data leakage to external clouds.

🎛️Fine Control

Flexible Speech Tuning

Full support for custom pitch, speed, volume, and multi-format output (MP3, WAV, AAC, OPUS, FLAC).

Supported Synthesis Engines

Choose the optimal engine according to your speed, hardware, and voice quality requirements.

EngineTypeVoices / LanguagesLatencyHardwareBest For
Edge-TTSRecommendedCloud / Free400+ Voices / 40+ Langs150 - 300msLow CPU / MemoryEveryday TTS, multi-lingual reading, audiobook production
Kokoro-82MUltra FastLocal ONNXEnglish / Japanese / Chinese< 150msCPU or CUDA GPUReal-time conversational voice agents, gaming NPC voiceover
CosyVoiceLocal NeuralZero-shot Voice Clone250 - 450msNVIDIA GPU (CUDA)High-fidelity personalized voice cloning, emotional speech
OpenAI RelayCloud Upstreamalloy, echo, fable, onyx, nova, shimmer300 - 600msAPI Key onlyOpenAI official fallback & multi-model proxying

Speech Synthesis Pipeline

From text request to audio stream in four streamlined processing steps.

Step 01

API Authentication & Parsing

Validates API key (Bearer/X-API-Key), parses parameters (model, voice, input, speed, format).

Step 02

Text Normalization & Cache Check

Computes query MD5 hash to return cached audio instantly, or segments text into natural phonetic chunks.

Step 03

Synthesis Engine Dispatch

Dispatches chunks to the designated engine (EdgeTTS, Kokoro, CosyVoice) with parallel synthesis.

Step 04

Audio Streaming & Output Cache

Encodes output format (mp3/wav/aac/opus) and streams chunks directly to client while caching result.

Zero SDK Changes Required

Seamlessly switch from OpenAI TTS to OneTTS by changing only base_url.

from openai import OpenAI

# Initialize client pointing to your local OneTTS gateway
client = OpenAI(
    base_url="http://localhost:8030/v1",
    api_key="sk-onetts-v1-k_8LxQm9Z2sTwE7rY1u"
)

# Synthesize speech stream using Edge-TTS or Kokoro
response = client.audio.speech.create(
    model="edgetts",
    voice="zh-CN-YunxiNeural",
    input="你好,这是通过 OneTTS 高性能语音合成网关生成的高保真音频流!",
    response_format="mp3",
    speed=1.0
)

# Stream or save audio directly
response.stream_to_file("output.mp3")

Released under the Apache-2.0 License.