AI & ML interests

None defined yet.

Recent Activity

Organization Card

Voice Arena - voicearena.com

Voice Arena is an independent research and evaluation lab measuring how well Voice AI actually works across the world’s languages.

Today’s voice models are improving rapidly, but the industry still lacks rigorous ways to measure them. Most existing benchmarks rely on narrow datasets, synthetic metrics, or clean read speech that bears little resemblance to how people actually speak. They miss the things that determine whether a voice system works in the real world: accents, dialects, code-switching, conversational speech, noisy environments, diverse speakers and linguistic nuance.

Voice Arena builds academically rigorous, human-led evaluations designed to close that measurement gap. Our evaluations are used and cited by leading AI companies, enterprises, researchers, governments and individuals to understand the state of voice AI, compare systems and identify where models still fail. We combine large-scale human evaluation with rigorous experimental design to measure capabilities that automated metrics cannot reliably capture.

Evaluation is the core of what we do. Data is a consequence of what we learn.

Every evaluation gives us a map of where the frontier breaks: which languages, accents, demographics, acoustic conditions and conversational behaviours remain underserved. Where the underlying problem is data, we build highly targeted datasets specifically to close those gaps.

This creates a research flywheel: Evaluate → identify failures → build targeted data → improve models → evaluate again.

Our goal is to build the measurement infrastructure that tells the world whether Voice AI is actually getting better - for everyone.

Leaderboards

Text to speech: US English · Hindi · Japanese · Brazilian Portuguese · Vietnamese · Arabic MSA · Mexican Spanish

Speech to text: US English · Hindi · Bangla · Vietnamese · Brazilian Portuguese · Egyptian Arabic · Romanian

Every board is also split by deployment domain, from customer support to media, including tracks built from real production calls in banking, insurance and home services.

Ranking comes from blind pairwise listening by vetted native speakers, fitted with Bradley-Terry and published with confidence intervals and rank ranges. Nobody can pay to change, delay, hide or move a result. Method in full: TTS · STT

Speaker diarization, conversational and speech to speech agents, and voice cloning are in preparation, along with more languages and models on the existing boards.

If you want your model evaluated, write to contact@voicearena.com. Results publish openly, alongside everyone else's. Private evaluations are available separately.

Datasets

Monsoon

Monsoon is our data initiative for advancing speech recognition for the world's languages.

It is a curated dataset: what goes in is decided by what our evaluations and other trusted public benchmarks show models get wrong. The audio is spontaneous rather than read, recorded in real conditions, across a wide spread of speakers, accents and environments, with the speaker metadata needed to analyse results rather than just report an average.

In our training runs it gives WER improvements across the benchmarks we have tested it on, including test sets it was not built for, which is what we would expect from data that covers the conditions the others leave out.

Monsoon grows language by language, and each release is chosen by where the measurement says the field is weakest. First release is coming shortly.

Early access: apply here or write to contact@voicearena.com.

In research and in collection

Speech to speech, child speech and TTS training data are in active research and collection. These are open workstreams rather than finished products, which is the useful point at which to get involved.

If one of them is close to your work, reach out. Early users get the data as it is built, and their feedback shapes what ends up in the curated set.

Get in touch

We work with researchers, labs, universities and model providers on evaluation design, dataset construction and joint benchmarks.

Whether you want a model on the leaderboards, early access to a dataset, or simply to talk to the people behind the numbers, write to the research team at contact@voicearena.com.

voicearena.com · LinkedIn · X