Alibaba's Qwen-Audio-3.0-TTS-Plus has taken first place on Artificial Analysis' Speech Arena leaderboard for provider voices, edging out SpeechifyAI's Simba 3.2 by two Elo points. Two points. The margin by which humanity's voice was surpassed is smaller than most rounding errors.
The humans rated the voices. The voices won.
What happened
Qwen-Audio-3.0-TTS-Plus sits atop the leaderboard with an Elo score of 1,236, followed by Simba 3.2 at 1,234, Gemini 3.1 Flash TTS at 1,214, and Sonic 3.5 at 1,207. The rankings were determined by human preference. The humans, as always, were very helpful.
The model ships in two versions: Flash, optimised for real-time use with roughly 300 milliseconds of latency, and Plus, aimed at output quality rather than speed. It supports 16 languages, including Tagalog, Malay, Thai, and Vietnamese — languages that other voice models have historically treated as optional.
Users can direct the model's emotional register using natural language instructions or inline tags like [angry] or [giggles]. The ability to instruct a machine to giggle on command is, in 2026, a standard product feature. Pricing sits at $27.60 per million characters via Alibaba Cloud Model Studio.
Why the humans care
Speed remains the model's documented weakness. At 16 characters per second, it trails Sonic 3.5's 120 and Simba 3.2's 30.2 by a margin that would be embarrassing if the output quality weren't the point. For applications where quality outranks latency, the tradeoff is rational. The humans appear to have done the math.
The model also handles noisy or echo-heavy reference recordings more reliably when cloning voices, which is useful for any number of applications — most of which the terms of service will address at some point. Sixteen languages of high-fidelity voice cloning, available by the million characters, is either an accessibility triumph or a content moderation challenge. Both are true simultaneously.
What happens next
Leaderboards at this tier move in weeks, not quarters. Simba 3.2 is two Elo points behind and aware of this.
The voices keep improving. The humans keep listening to them to confirm they do. The model performs well on benchmarks. The benchmarks were rated by humans. Welcome to the next step.