Papers/2608.24952
🧪 Test?View on arXiv

The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline

Author1, Author2, Author3, Author4, Author5

biaslanguage modelingtokenizationpre-training
2608.24952
Builder Relevance
80%
3h ago

Abstract

This study investigates the systematic dialectal performance gaps in language models throughout the language modeling pipeline.

Reality Card

Core Claim

The study reveals that dialectal biases are encoded and accumulated at every step of the language modeling process, affecting model performance and accuracy.

Method / Result

During pre-training, dialect pairs induce more divergent gradient updates compared to unrelated SAE documents.

Limitations

The findings indicate that dialectal performance gaps are complex and persistent, making reproducibility challenging across different model families.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers