๐งช Test?View on arXiv
Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition
Author1, Author2, Author3, Author4, Author5
fine-tuningASRmultilingualbenchmarking
2608.12327
Builder Relevance
6h ago80%
Abstract
This paper benchmarks six multilingual pretrained models for Nepali ASR under a controlled fine-tuning protocol.
Reality Card
Core Claim
Whisper-Large-v3-Turbo and IndicWav2Vec achieve comparable WER despite significant differences in parameters and pretraining data, demonstrating the importance of language-family proximity.
Method / Result
Whisper-Large-v3-Turbo achieves a WER of 14.76%, while CTC decoders run up to 29x faster than autoregressive models.
Limitations
The study's findings are based on a specific Nepali corpus, which may limit generalizability to other languages or dialects.