Alibaba rolled out Qwen Audio 3.0 TTS, a new hosted text to speech model built to deliver very natural and expressive AI voices in several languages. This launch gives developers and businesses an extra tool if they want to power voice driven apps, virtual assistants, content creation platforms, or customer service products. With support for multiple languages, voice customization, and low latency performance, the model is designed to cover a wide range of speech AI needs.
It also shows Alibaba’s continued push into the speech AI space. Qwen Audio 3.0 TTS comes in several service tiers, so users can pick between faster responses or higher quality sound. Along with its multilingual reach, the model now lets users directly guide the emotion, style, and delivery of AI voices with simple instructions. That flexibility makes the platform a strong choice for developers who want control but do not want to work with complex voice editing.
What Qwen Audio 3.0 TTS Can Do
Qwen Audio 3.0 TTS at a glance

Qwen Audio 3.0 TTS is a hosted text to speech model from Alibaba’s Tongyi Lab that converts written text into realistic and expressive speech. It supports a variety of use cases ranging from AI assistants and educational tools to customer support systems and digital content creation.
Introducing the Qwen-Audio-3.0-TTS.
— Qwen (@Alibaba_Qwen) July 23, 2026
Our latest text-to-speech model, in two flavors:
• Flash: real-time interaction
• Plus: high-quality generation
What's new:
• Fine-grained inline tags-steer [whisper], [angry], [breaths] & [laughs]
• Free-style natural-language… pic.twitter.com/XZMNeK0lFl
Some of the key features announced with the release include:
- Support for 16 languages, allowing developers to generate speech for multilingual applications.
- Availability in Flash and Plus service tiers, giving users a choice between lower latency and higher quality voice generation.
- Natural language voice instructions, making it possible to describe how the generated speech should sound instead of relying on complex settings.
- Support for 86 inline control tags that can adjust speech characteristics such as emotion, laughter, whispers, pauses, and other speaking styles directly within the text.
- Voice cloning capabilities that enable users to create speech using a voice similar to a provided reference.
- Expressive speech generation with the ability to produce different tones and delivery styles for a more natural listening experience.
- Hosted deployment, allowing developers to access the model through Alibaba Cloud without managing the underlying infrastructure.
- Support for both streaming and non streaming speech generation, depending on application requirements.
- Voice design options that give users additional flexibility when creating AI generated voices.
- API based access for easier integration into applications and services.
According to the information shared during the launch, the combination of voice instructions and inline control tags gives users more direct control over speech output. Instead of manually adjusting technical voice parameters, developers can describe the intended speaking style or insert control tags within the text. The multilingual support also makes the model suitable for applications that need to serve users across different regions. By supporting 16 languages, developers can build products that generate speech for a broader audience while using the same hosted platform. Another notable addition is the voice cloning feature. This allows developers to generate speech based on a reference voice, expanding the model’s use in personalized assistants, narration, and interactive voice applications.

How to Access Qwen Audio 3.0 TTS on Alibaba Cloud
Alibaba released Qwen Audio 3.0 TTS as both Flash and Plus service tiers. Flash focuses speed, while Plus focuses on the speech quality, so one can pick what fits their project best. Interested users can find the hosted model in Alibaba Cloud Model Studio, where developers can access the text to speech API and plug it into their software. The hosted approach removes the need for users to deploy and maintain their own speech generation infrastructure.
The release information also highlights several technical capabilities available through the API, including streaming and non streaming speech generation, voice cloning, voice design, and instruction based speech control. Performance wise, the launch highlights the model’s multilingual ability, its expressive voices, and the detailed speech customization you get from those 86 control tags and natural language prompts. Altogether, these features make Qwen Audio 3.0 TTS a flexible tool for anyone creating AI powered voice apps.
Also read: Alibaba Unveils Qwen-Image-3.0 With Richer Detail and Deep Knowledge









