The scoreboard gap in voice AI

Voice AI has quietly become one of the fastest-growing corners of open machine learning. The Hugging Face Hub listed more than 8,000 text-to-speech models as of September 30, 2026, according to the platform's own announcement that day. Yet the rankings most developers consult told a very different story about which voices matter. That mismatch is what a new public resource, the Open TTS Leaderboard, was built to address.

The leaderboard, launched on September 30, 2026, by Hugging Face audio researchers Eric Bezzam, Steven Zheng and Eustache Le Bihan together with community contributor mrfakename, evaluates open-source speech models using objective, reproducible measurements. Its stated purpose is to give freely available models a fair, fast test against each other, particularly across languages beyond English.

For nontechnical readers, the shift matters because it changes how anyone can pick a synthetic voice. Instead of relying on a popularity contest dominated by commercial vendors, builders of accessibility tools, language-learning apps and voice agents can now compare open models using measurable error rates, latency and voice-cloning fidelity.

Why arenas favor paid APIs

The most widely referenced TTS rankings today are vote-based arenas. TTS Arena v2 and the Artificial Analysis Voice Arena present users with two generated clips and ask which sounds better, then compute an Elo-style score from the votes. Human preference is, as the Hugging Face authors put it, the gold standard: in the end, what matters is what listeners actually prefer.

But this method has two structural problems the authors identify. First, throughput. As the blog states, "arenas cannot scale to keep up with the pace of TTS releases." Collecting enough votes to rank a single model can take weeks, while thousands of models appear on the Hub.

Second, representation. The authors report that "as of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena." They attribute this to practical factors: adding a commercial API model requires little more than an API key, while an open model must be hosted and served by the arena operator, and commercial providers have stronger incentives to seek placement.

The result is that most of the 8,000-plus open TTS models had no reliable public comparison at all. A developer wanting a free, locally runnable voice model for, say, Spanish or Korean had no standardized way to know which candidates actually worked.

Three measurable dimensions instead of votes

The Open TTS Leaderboard replaces votes with three families of objective measurements, all computed on the same hardware.

Intelligibility is measured by transcribing each model's generated audio with Qwen3 ASR, described as the top-ranking open-source model on Hugging Face's Open ASR Leaderboard, and comparing that transcript to the original text. The word error rate (WER) or character error rate (CER) between the two gives a concrete proxy for how understandable the speech is. English models are ranked on macro-average WER over the English splits of two public benchmarks, Seed TTS Eval and CV3 Eval (zero shot).

Speed is measured two ways: inverse real-time factor (RTFx) for batched offline inference, and time-to-first-audio (TTFA) for streaming latency, both on an NVIDIA H200 GPU, with CPU results available for a growing subset of models. TTFA matters for interactive applications like voice agents, where users cannot tolerate a long silent pause after speaking.

Speaker similarity (SIM) uses cosine similarity between WavLM speaker embeddings of the generated audio and a reference clip. This is the voice-cloning measure: how closely the model preserves the identity of a voice it was given as a reference.

The practical effect is speed of evaluation. According to the blog, objective measurement reduces evaluation time "from a couple weeks (for collecting votes) to a couple hours." That pace makes it feasible to evaluate new models as they ship, rather than months later, if ever.

English scores are not the whole world

A central finding of the launch is that English rankings are a poor guide elsewhere. The authors state plainly in the blog that "English performance doesn't necessarily translate to other languages," which is why multilingual ranking is a first-class feature rather than an afterthought.

On the default English view, the authors name hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro as leading on English WER averaged across the two benchmark splits. When users toggle to multilingual views, they report that k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are strong multilingual models.

These named leaders are the authors' summaries of their own rankings as of September 30, 2026, not independent verification by this publication; leaderboards are living pages and rankings will shift. But the pattern is notable: two of the three English leaders also appear in the multilingual tier, and several of the strongest open models are ones a vote-based arena might never have gotten to.

The leaderboard also surfaces a subtle technical point: some models, the authors note, including bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, improve under voice cloning, that is, when a reference audio sample is provided. For communities whose languages are poorly served by mainstream voice products, this distinction between zero-shot and cloning performance is exactly the kind of information that was previously unavailable.

What the leaderboard cannot tell you

The authors are explicit about the limits, and it is worth quoting their wording in full: "Neither directly measures naturalness, expressiveness, or listener preference."

In other words, a model can produce highly intelligible speech that still sounds flat, robotic or unpleasant. ASR-based WER is a proxy for intelligibility, and speaker similarity estimates identity preservation; neither tells you whether a human listener would enjoy the result. The leaderboard is designed to complement human preference arenas, not replace them. Its Pareto plots, which visualize tradeoffs between error rate, speed and model size, can even help arena operators decide which models to include in human evaluations.

Other open questions the authors acknowledge: votes collected on the "Listen" tab are not yet included in rankings, though they may be later as volume accumulates and spam is filtered via HF account login. And the evaluation scripts are not yet open-sourced; the team says it "will soon open-source the evaluation scripts," following the pattern of the Open ASR Leaderboard repository, which will allow community scrutiny and contributions through GitHub issues and pull requests.

A final caveat on fairness: hardware is standardized to H200 GPUs and CPU runs, but model defaults such as voice choice and sampling configuration still reflect each model's own settings, and benchmark coverage is uneven, with Seed TTS Eval containing audio only for English and Chinese.

Why this matters beyond the rankings

For most readers, the significance is straightforward. The free-to-run side of voice AI finally has a scoreboard that measures what it claims to measure, in hours rather than weeks.

Three concrete consequences follow. First, developers building on modest budgets can choose open voice models on evidence, comparing error rates and latency directly instead of guessing or trusting vendor marketing. Second, accessibility builders, who often need dependable, cheap, locally runnable speech, gain a practical screening tool; a screen reader or communication aid cannot afford models that mangle words, and WER now makes that visible. Third, speakers of languages other than English, historically the most underserved users of voice technology, can see which multilingual open models genuinely work in their language rather than being offered a repackaged English-centric ranking.

There is also an ecosystem effect. When evaluation is cheap and open, neglect stops being self-reinforcing. Models that were invisible because nobody could afford to host them for votes can now be measured and, if good, discovered. That does not guarantee quality, and objective metrics will miss what human ears catch. But it removes the structural excuse that open models were too numerous and too awkward to evaluate at all. Whether the community picks up the evaluation scripts when they are released, and keeps the rankings honest, will determine whether this becomes durable infrastructure or another well-intentioned launch.