Qwen-Audio-3.0-TTS: More Multilingual, Easier to Direct
Our text-to-speech model, now across 16 languages.
Qwen-Audio-3.0-TTS is our latest text-to-speech model release. It ships as two variants from the same lineage:
Flash: tuned for real-time interaction, with a first-packet latency at 300ms-level.
Plus: tuned for high-quality generation, where naturalness and timbre fidelity matter more than speed.
API: Model Studio
Release Blog: Blog
This release focuses on four things developers actually run into in production: broader language coverage, natural-language style control, fine-grained tag control, and robustness when the reference audio isn’t clean.
Qwen-Audio-3.0-TTS-Plus currently ranks #1 on Artificial Analysis, the independent third-party TTS leaderboard.
Here’s what changed and what the numbers look like.
Multilingual Coverage Across 16 Languages
Qwen-Audio-3.0-TTS was optimized across English, Chinese, Japanese, Korean, German, and 16 languages total, plus improved fidelity on several Chinese dialects.
Supported languages (16): Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, Vietnamese.
*Full language support rolling out soon.
WER/CER (lower is better): Qwen-Audio-3.0-TTS family demonstrates the strongest overall multilingual intelligibility, achieving the best WER/CER in 10 of 16 languages. Flash delivers the lowest average WER/CER at 3.87, while Plus remains highly competitive at 3.96, both outperforming the other systems on average.
Speaker Similarity (higher is better): Qwen-Audio-3.0-TTS shows a clear and consistent advantage in speaker similarity. Plus ranks first across all 16 languages with an average SS of 82.75, while Flash follows at 80.44, demonstrating strong and robust voice-preservation quality across diverse languages.
Demos
Reference audio:
Outputs:
Arabic
أعلنت الحكومة اليوم عن خطة جديدة لتطوير البنية التحتية في المناطق الريفية بتكلفة تقدر بمليارات الدولارات.
Indonesian
Sekolah dasar membentuk karakter anak melalui pendidikan akademik dan moral sejak dini.
Portuguese
No programa de hoje, vamos descobrir os sabores autênticos da culinária alentejana, conhecida por seus pratos à base de pão, ervas aromáticas e azeite de oliva extra virgem.
Thai
เราขออภัยสำหรับความล่าช้าในการจัดส่งคำสั่งซื้อของคุณ เนื่องจากสภาพอากาศที่ไม่ดีซึ่งส่งผลกระทบต่อระบบโลจิสติกส์ของเรา
Vietnamese
Sức khỏe thể chất và tinh thần luôn song hành với nhau nên ta phải chăm sóc cả hai một cách cân bằng và điều độ.
Malay
Perbicaraan juga ditanda tarikh untuk suspek menjalani perbicaraan dengan cepat.
Filipino
Katumbas ng kalahating oras ang pamamasyal sa makakatawag ng pansin na lugar.
Chinese dialect (Hangzhou)
爬北高峰爬得我吃力煞,腿都抖抖战了。好在顶上风吹吹,毛落胃的。
Chinese dialect (Shannxi)
额天天跑得气喘弹腾的,你还吃呢?赶紧把嘴管住,嫑让那肉在腰上堆着咧。
Chinese dialect (Shanghai)
屋里向装修真个是烦人,天天敲墙壁,声音响得来!灰尘落得一塌糊涂,真想快点弄清爽。
Style Control In Natural Language
You can describe the delivery you want in natural language instead of hand-tuning acoustic parameters.
Simple prompts with plain language steer emotion, role, scenario, and pace without any labeling expertise.
Demos
Prompt: Say the following angrily.
Target text:After the meeting, he organized the materials into a folder and put the laptop back in his bag.
Reference audio:
Generated audio:
Prompt: Large hall projection, broad pacing, slight reverberant feel, lifted intonation on the welcome — a stadium announcer welcoming the crowd.
Target text:Good evening, everyone, and welcome back to the stadium tonight. Conditions are perfect for a great game.
Reference audio:
Generated audio:
Fine-Grained Tags For Non-Verbal Details
When you need precise control over the non-verbal details — a breath, a laugh, a shift in tone — you can embed inline tags directly in the target text, like [gasp], [giggles], or [angry].
This makes the model useful for narration, games, and dubbing where the non-verbal cues carry as much as the words.
Demos
Target text:[angry] I told you three times to lock the side door, and now the entire shipment is sitting in the rain.
Reference audio:
Generated audio:
Target text:[giggles] I may have replaced every office pen with one that writes in glitter, but I refuse to call it a crime.
Reference audio:
Generated audio:
More Robust Voice Cloning From Imperfect Audio
Reference clips from the real world are rarely studio-clean. Qwen-Audio-3.0-TTS was trained with targeted acoustic simulation so speech enhancement is built into the cloning path. The model suppresses reverb and noise while preserving timbre.
In our high-noise and high-reverb tests, this release produced noticeably cleaner output than previous versions from the same degraded references.
Demos
High-Noise Inputs:
Target text:Please keep the shared kitchen clean by washing your dishes immediately after use. Food left in the refrigerator over the weekend will be discarded every Monday morning.
Reference audio:
Generated audio:
Qwen-Audio-3.0-TTS:
Previous version:
High-Reverb Inputs:
Target text:The coffee shop was quiet except for the gentle hum of the machine.
Reference audio:
Generated audio:
Qwen-Audio-3.0-TTS:
Previous version:
Also In This Release
A curated preset voice library spanning 16 supported languages, so you can ship a voice without cloning one first.
48 kHz audio output (coming soon).
Try It
Qwen-Audio-3.0-TTS is available now. Grab the model here:
API: Model Studio
Release Blog: Blog
If you build something with it, we’d genuinely like to hear what worked and what didn’t — the failure cases are where the next version comes from.





it would be great if add Persian language support in next version