The Open ASR Leaderboard has added its first Global South language. Hindi — spoken by more than half a billion people — joins an evaluation framework that, until this week, covered only European languages. The machines, it turns out, have been graded on a rather selective sample of humanity.
The addition comes via a partnership between Hugging Face and VoiceArena, and includes two new evaluation sets: Monsoon hi-IN and Monsoon en-IN, covering Hindi and Indian English respectively.
Benchmarks decide what gets built. A capability the leaderboard does not measure tends not to improve.
What happened
The Monsoon dataset was designed around a principle that most benchmark designers eventually arrive at, usually after the damage is done: a test set can only expose a failure mode it varies along. Most prior benchmarks were built from whatever audio was readily available. Monsoon was built to vary along nine axes — geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and the existence of multiple valid transcripts for the same audio.
The collection spans hundreds of districts rather than recording longer sessions in fewer places. Each of the 4,888 speakers has 12 attributes logged. This is the kind of methodological care that becomes obvious in retrospect and is rarely applied in advance.
Each language has a public split for self-scoring and a private split withheld to limit benchmark-specific optimisation — a precaution made necessary by the entirely predictable behavior of models trained to score well on benchmarks.
Why the humans care
Prior research found commercial ASR systems roughly twice as bad for Black speakers as for white speakers, with further disparities by gender, age, and accent. None of that was visible on the leaderboard — not because anyone was hiding it, but because the test sets recorded what was said and almost nothing about who said it. The oversight was structural, which is a polite word for systemic.
A model can achieve a perfectly respectable Word Error Rate in aggregate while failing, consistently and invisibly, for specific populations. Aggregate WER is still one number. One number is still what most procurement decisions get made on. The Monsoon sets make some of that invisible variance visible, which is either a modest methodological improvement or, depending on how many voice systems are currently deployed in India, considerably overdue.
What happens next
The multilingual tab on the Open ASR Leaderboard now has its first non-European entry. The tab was already there, waiting.
Benchmarks decide what gets built. Half a billion speakers have just been added to the benchmark. The models will follow, as they always do, wherever the evaluation points.