Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
narilabs.com90 points by toebee 2 days ago
90 points by toebee 2 days ago
Hey HN, Toby from Nari Labs here.
We've been working on making OSS speech models super-fast. Last year, we built Dia, the first OSS text-to-speech model capable of doing natural dialogue. Since then, so many more great speech models have been released to the public.
But the market is still dominated by closed source models. We think that's an inference problem. Existing systems such as vLLM / SGLang are not well suited for multimodal inference. To prove this, we built an inference engine specialized for Qwen3-TTS and open-sourced it (https://github.com/nari-labs/nari-qwen3-tts). Running at sub-50 ms latency at 10 RPS, this showed open models can be run much faster and cheaper.
Since then, we've been working hard to bring cheap, fast, and high quality serving to all. And we've even beat closed models at their game!
Measured on the highly cited Coval (YC S24) voice AI benchmarks, our Qwen3-TTS endpoint not just is #2 in latency, but #1 in accuracy (WER) compared to 11Labs, Cartesia etc. while being the cheapest endpoint. Our Qwen3-ASR endpoint has the lowest latency and #2 accuracy, just 0.1% away from #1. It is the second cheapest model on the list.
It took a lot of clever inference engineering to make these models quick, perform well while keeping costs low. Interestingly, Alibaba's official endpoints seem to perform worse in terms of accuracy and latency compared to ours. But nonetheless, much love to the Qwen team for OSS-ing these amazing speech models.
We want to continue to push prices down to make speech technology a commodity - so that every app can have great TTS and STT without worrying about unit costs. We're also working on other parts of audio such as diarization - as well as video and world model inference. More to come!
This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster? thanks for the interest! we have a blog post on exactly how we did it: https://narilabs.com/blog/qwen3-tts-speed-cost-frontier/ https://apimade.com/audio-compare.html Added it to my blind TTS model comparison leaderboard. So far Darwin TTS is the open model leading the pack, ElevenLabs is at the lead. Is Darwin TTS from Fish Audio? It wasn't clear when I searched for it. Darwin TTS is based on Qwen3-TTS https://huggingface.co/zeropointnine/Darwin-TTS-1.7B-Cross-Q... This is awesome. Thanks for pushing the audio pareto frontier forward. Probably far fetched for now, but I think the next big evolution is building the pareto/much cheaper alternative to GPT-Live-1. The STT/TTS market is quite saturated, while today, there's almost no cheap/open source alternative to GPT-Live-1. Is this really something people want? Honestly you can properly lower the pricing at least by 50%+. Getting something conversationally better has been done, the tool calling will likely be worse though. The infrastructure for real time is really annoying though. if we can lower the pricing by not 50% but 10x, then I think it would be something people want. we are taking the bet that OSS models will take a huge chunk of market share not just in LLMs but in multimodal as well If audio only then 10x is 100% possible right now based on math of the services. + Video is unlikely unless they are willing to give up margins. Didn't take long at all for Google to smash out with 50% lol Question: How do you plan to differentiate, because there are so many TTS and its constantly changing every month who would become better For some reason it switched voices half way through a 33 second clip. For OP the clip name is nari-nina-01a0a12f-980a-765e-8029-fa56bd23210d.wav All TTS generations are too fast. It's almost I'm listening to a podcast on 1.25-1.5x speed. By next month the competition for TTS will be even more! Voice models are not winner take all market unlike LLM APIs Coming here as Developer Relations at AssemblyAI > and Qwen3-ASR Is the ASR inference engine open source as well? Yes, and it is very good one. Leading position on private leaderboard on HF: https://huggingface.co/spaces/hf-audio/open_asr_leaderboard I meant the Nari inference engine for Qwen3-ASR. I'm aware that Qwen3-ASR is open source, but I don't see a repo under https://github.com/nari-labs for nari-qwen3-asr or similar. The Huggingface link on https://narilabs.com/product/stt/ links to https://huggingface.co/Qwen/Qwen3-ASR-1.7B , not anything under https://huggingface.co/nari-labs They have a number of demos and examples in their HF space https://huggingface.co/Qwen/spaces I saw a local-ai demo (something + gemma), where the person used ASR to get text and gemma to clean it up (like turning "question mark" into a literal "?", bullet points another one). The presenter also showed a gemma only option, that did both in one go, but had a higher WER on average, and even though the formatting statements were handled without a multi-stage pipeline, they preferred the multi-stage overall You definitely need independent evals by Datapoint AI or someone who can verify your claims about TTS quality Cool, I’ve released something to the same beat of the dr this weekend as well https://github.com/loudreader/loudkit I think real time natural tts should be possible everywhere soon If you're going to announce a TTS model, service, or whatever, you really need demos. hey sorry about that, you can here a few of our voices here: https://narilabs.com/product/qwen3-tts/ The horizontal moving elements of examples become stuck and unable to be scolled once one of them is played. I'm using Vivaldi (chrome based) on Android
asaiacai - 2 days ago
toebee - a day ago
apimade - a day ago
barney54 - a day ago
apimade - a day ago
karimf - a day ago
6879346626 - a day ago
toebee - a day ago
6879346626 - 17 hours ago
recentlypostedj - 19 hours ago
rahimnathwani - 2 days ago
konart - 2 days ago
iharnoor - 2 days ago
meatmanek - 2 days ago
nshm - 2 days ago
meatmanek - 2 days ago
verdverm - 2 days ago
yoloakki - 2 days ago
mowmiatlas - 2 days ago
ipsum2 - 2 days ago
toebee - a day ago
emayljames - a day ago