Overview & Architecture
OneTTS is an enterprise-grade, high-performance Text-to-Speech (TTS) aggregation gateway designed to unify disparate speech synthesis backends into a single, OpenAI-compatible API endpoint.
🌟 Why OneTTS?
- OpenAI Standard Drop-in Replacement: Any application supporting OpenAI TTS (e.g. Dify, FastGPT, Open WebUI, DuRT) can switch to OneTTS simply by updating
baseURLandapiKey. - Zero-GPU Startup: With built-in Edge-TTS support, deploy immediately on low-cost VPS or local CPU machines with zero GPU requirements.
- Low Latency Streaming: Pipeline sentence chunking delivers first-chunk audio in less than 200ms.
- Intelligent Audio Caching: Duplicate synthesis requests hit the local MD5 cache instantly, saving bandwidth and compute.
🏗️ Architecture
text
Client (DuRT / Dify / Web / API)
│
▼ (HTTP /v1/audio/speech)
┌───────────────────────────────┐
│ OneTTS Gateway │
│ ┌─────────────────────────┐ │
│ │ MD5 Synthesis Cache │ │
│ └─────────────────────────┘ │
│ ┌─────────────────────────┐ │
│ │ Sentence Text Splitter │ │
│ └─────────────────────────┘ │
└──────────────┬────────────────┘
│
┌────────────┼────────────┐
▼ ▼ ▼
┌─────────┐ ┌─────────┐ ┌───────────┐
│ EdgeTTS │ │ Kokoro │ │ CosyVoice │
└─────────┘ └─────────┘ └───────────┘