Skip to content

Overview & Architecture ​

OneTTS is an enterprise-grade, high-performance Text-to-Speech (TTS) aggregation gateway designed to unify disparate speech synthesis backends into a single, OpenAI-compatible API endpoint.


🌟 Why OneTTS? ​

  • OpenAI Standard Drop-in Replacement: Any application supporting OpenAI TTS (e.g. Dify, FastGPT, Open WebUI, DuRT) can switch to OneTTS simply by updating baseURL and apiKey.
  • Zero-GPU Startup: With built-in Edge-TTS support, deploy immediately on low-cost VPS or local CPU machines with zero GPU requirements.
  • Low Latency Streaming: Pipeline sentence chunking delivers first-chunk audio in less than 200ms.
  • Intelligent Audio Caching: Duplicate synthesis requests hit the local MD5 cache instantly, saving bandwidth and compute.

🏗️ Architecture ​

text
 Client (DuRT / Dify / Web / API)
                │
                ▼ (HTTP /v1/audio/speech)
   ┌───────────────────────────────┐
   │        OneTTS Gateway         │
   │  ┌─────────────────────────┐  │
   │  │  MD5 Synthesis Cache    │  │
   │  └─────────────────────────┘  │
   │  ┌─────────────────────────┐  │
   │  │ Sentence Text Splitter  │  │
   │  └─────────────────────────┘  │
   └──────────────┬────────────────┘
                  │
     ┌────────────┼────────────┐
     ▼            ▼            ▼
 ┌─────────┐ ┌─────────┐ ┌───────────┐
 │ EdgeTTS │ │ Kokoro  │ │ CosyVoice │
 └─────────┘ └─────────┘ └───────────┘

Released under the Apache-2.0 License.