BuzzASR Boosts Low-Resource Speech Accuracy
New model suite beats Whisper in 77 of 102 languages, cutting error rates by 2.8x.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
The BuzzASR model suite outperforms Whisper-large-v3 in 77 of 102 languages, reducing the average Character Error Rate (CER) by more than 2.8 times.
Mainstream end-to-end speech recognition models are typically trained on multiple languages, which often leads to poor performance in low-resource languages with limited training data. While monolingual fine-tuning is known to be effective, it had previously been applied only to a small number of languages.
This study scales the monolingual fine-tuning strategy to 102 languages and introduces tokenizer replacement, improving compression rates by an average of 3.3 times. All models, code, and detailed results have been open-sourced.