🧪 Test?View on arXiv
The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline
Author1, Author2, Author3, Author4, Author5
biaslanguage modelingtokenizationpre-training
2608.24952
Builder Relevance
3h ago80%
Abstract
This study investigates the systematic dialectal performance gaps in language models throughout the language modeling pipeline.
Reality Card
Core Claim
The study reveals that dialectal biases are encoded and accumulated at every step of the language modeling process, affecting model performance and accuracy.
Method / Result
During pre-training, dialect pairs induce more divergent gradient updates compared to unrelated SAE documents.
Limitations
The findings indicate that dialectal performance gaps are complex and persistent, making reproducibility challenging across different model families.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.