September 23, 2026
Qwen-Audio Cuts Audio Transcription Prices: ASR Is Up to 95% Cheaper
On September 23, Alibaba cut Qwen-Audio prices: TTS by about 70%, Realtime by about 85%, and ASR by up to 95%. The lineup includes five models for transcription, synthesis, real-time dialogue, and audio generation. The new TTS-Next and ASR-Next generate voice and sounds or label speakers and timestamps in transcripts.

Qwen
@alibaba_qwen
⚡ Meet Qwen-Audio-3.1! ASR, TTS, and Realtime have all been fully upgraded, with two new models joining them: TTS-Next for audio creation and ASR-Next for understanding it. Five models, one complete audio stack: understanding, generation, interaction, and creation. Plus major price cuts across the lineup: TTS by about 70%, Realtime by about 85%, and ASR by up to 95%. Highlights: 🥳 - ASR: better recognition across languages and dialects, while built-in cleanup automatically removes filler words and repetitions. Transcripts become cleaner and more coherent. - ASR-Next: supports multi-speaker transcription with labels, timestamps, and synchronized text. Understands emotions, background sounds, and machine sounds for audio descriptions, event detection, audio verification, and reasoning about audio. - TTS: synthesizes speech across languages and dialects with natural cross-lingual voice transfer. Emotion, speed, and style can be set with simple instructions. - TTS-Next: a unified single-pass LM and diffusion architecture creates voices, sound effects, and background audio for audiobooks, podcasts, games, and advertising. - Realtime: you can speak and listen at the same time, and interrupt the model at any moment, like in a normal conversation. When it detects a bad mood, it slows down and responds with empathy. Unlock the full potential of Qwen-Audio-3.1! 👇 - Blog: https://fun-resource-shanghai.oss-cn-shanghai.aliyuncs.com/cuijiayan.cjy/tmp/exp/qwen_audio_3_tts_blog_review_260918/index.shtml?Expires=2105366399&OSSAccessKeyId=LTAI5tQrCBwj82sVMCWoSmzE&Signature=9ZaZIEahoN41Zb9MFJjf34G%2BSCM%3D - Qwen-Audio-3.1-ASR: https://www.qwencloud.com/models/qwen-audio-3.1-asr-flash-filetrans - Qwen-Audio-3.1-Realtime: https://www.qwencloud.com/models/qwen-audio-3.1-realtime-plus - Other APIs: coming soon from @qwen_cloud

· 160.3K views
Until September 22, Qwen-Audio-3.0-Realtime-Flash cost $4.5 per 1 million audio input tokens and $15 for text+audio output in Singapore. From September 22, Alibaba charges $0.93 and $1.87, and on September 23 it announced discounts for the Qwen-Audio-3.1 lineup.
The implementation details are ready. The file transcription model accepts recordings up to 12 hours or 2 GB, separates speakers, and works over HTTP. The streaming variant connects over WebSocket and has no duration limit. TTS-Next accepts up to three 30-second reference clips and generates up to 240 seconds of podcast audio per request.
Alibaba says it will add other Qwen-Audio-3.1 APIs later.
