Hacker News
How We Made a Text-to-Speech Model Respond in Sub-50 ms
The authors implement Qwen3-TTS CustomVoice 1.7B on a single NVIDIA H100 SXM, achieving 10 RPS with sub-50 ms p95 time-to-first-audio while keeping underruns at zero and maintaining intelligible speech. After tuning latency knobs and trimming leading silence, their solution outperforms vLLM-Omni, SGLang-Omni△, VoxServe, and M* across Poisson open-loop traffic, sustaining sub-50 ms TTFA up to 10 RPS and under 100 ms at 20 RPS.