Tech & AI News
Hacker News

Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

Nari Labs released an inference engine for Qwen3-TTS that runs at sub-50 ms latency with 10 RPS, proving open models can be fast and cheap. On the Coval voice-AI benchmark, Qwen3-TTS ranks #2 in latency, #1 in accuracy (WER) against 11Labs and Cartesia, and is the cheapest endpoint. Qwen3-ASR achieves the lowest latency, #2 accuracy only 0.1% behind the leader, and is the second-cheapest model, outperforming Alibaba’s official endpoints in both accuracy and speed.