Papers/2609.38219
πŸ“– Read?View on arXiv

TutlAit v1: a crowdsourced Moroccan Tamazight speech dataset with Arabic transcriptions and regional accent labels

Author1, Author2, Author3, Author4, Author5

speech recognitioncrowdsourcingdataset creationaccent identification
2609.38219
Builder Relevance
80%
1h ago

Abstract

The TutlAit dataset provides a significant resource for Moroccan Tamazight speech technology by offering labeled audio paired with Arabic transcriptions and regional accent labels.

Reality Card

Core Claim

The TutlAit dataset comprises 13,384 audio files of Moroccan Tamazight speech, enabling advancements in speech recognition, translation, and accent identification.

Method / Result

The dataset includes 75,231 seconds of audio, with the Atlas variety accounting for 9,956 files.

Limitations

The dataset's reliance on crowdsourced contributions may lead to variability in transcription quality and regional representation.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers