π Read?View on arXiv
TutlAit v1: a crowdsourced Moroccan Tamazight speech dataset with Arabic transcriptions and regional accent labels
Author1, Author2, Author3, Author4, Author5
speech recognitioncrowdsourcingdataset creationaccent identification
2609.38219
Builder Relevance
1h ago80%
Abstract
The TutlAit dataset provides a significant resource for Moroccan Tamazight speech technology by offering labeled audio paired with Arabic transcriptions and regional accent labels.
Reality Card
Core Claim
The TutlAit dataset comprises 13,384 audio files of Moroccan Tamazight speech, enabling advancements in speech recognition, translation, and accent identification.
Method / Result
The dataset includes 75,231 seconds of audio, with the Atlas variety accounting for 9,956 files.
Limitations
The dataset's reliance on crowdsourced contributions may lead to variability in transcription quality and regional representation.
Paper to code
Verified implementation resources so builders can test the paperβs claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.