AI
Alibaba Cut Voice AI Prices Up to 95 Percent, and Every App on Earth Just Gained a Voice
Alibaba just made voice the cheapest sense a machine can have. The Qwen team released Qwen-Audio-3.1 on September 23, a five model stack covering speech recognition, text to speech, and live conversation, and paired it with price cuts that reshape the economics of every voice product. Text to speech drops about 70 percent, live conversation roughly 85 percent, and speech recognition as much as 95 percent.
Start with the ears. The new ASR model handles multiple languages and regional dialects, and it cleans transcripts as it goes, stripping filler words like um and uh plus repeated phrases, so the output reads like something a person meant to say. ASR-Next goes further, adding speaker labels with timestamps, emotion detection, and awareness of ambient sounds and machine noise. A raw call recording becomes a structured, searchable record of who said what and when.
Then the voice. The TTS model speaks multiple languages and carries a voice across languages, so a voice cloned from an English recording can deliver the same lines in another tongue. Builders steer emotion, pace, and style with plain instructions, like asking for a sharp, commanding tone that demands respect. TTS-Next folds a language model together with a diffusion audio generator to produce voice, sound effects, and background audio in one pass, the kind of single step production that once needed a studio.
The fifth model, Realtime, holds both sides of a conversation at once, listening while it speaks and letting a person interrupt mid sentence. It also reads the room. When it detects a low mood in a speaker, it slows down and answers with more empathy. One model in the lineup supports 30 languages and 16 Chinese dialects, with first word streaming responses landing around 160 milliseconds.
Here is the angle that matters. A 95 percent price cut is Alibaba declaring that speech is infrastructure now, a utility like bandwidth rather than a premium feature. Workflows that once needed a serious budget, transcribing every customer call, dubbing video into a dozen languages, translating conversations live, now pencil out for small teams and solo builders. The moat moves from owning the model to owning the distribution, the workflow, and the trust of the user.
The practical read is simple. Voice interfaces are about to show up in far more products than the flagship assistants, because the price floor just fell out from under them. Builders who were waiting for the price to make sense can stop waiting. Everyone else gets to talk to their software sooner than expected, and in more languages than they speak.
Quick answers
What is this story about?
Alibaba just made voice the cheapest sense a machine can have. The Qwen team released Qwen-Audio-3.1 on September 23, a five model stack covering speech recognition, text to speech, and live conversation, and paired it with price cuts that reshape the economics of every voice product. Text to speech drops about 70 percent, live conversation roughly 85 percent, and speech recognition as much as 95 percent.
Why does this story matter?
The practical read is simple. Voice interfaces are about to show up in far more products than the flagship assistants, because the price floor just fell out from under them. Builders who were waiting for the price to make sense can stop waiting. Everyone else gets to talk to their software sooner than expected, and in more languages than they speak.
Sources
New to crypto? Read the crypto glossary, browse frequent questions, read our story, or explore the story archive.