Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias
Author1, Author2, Author3, Author4, Author5
Abstract
This study reveals that language models exhibit representational biases regarding user competence, even when behavioral biases are not apparent.
Reality Card
The study demonstrates that demographic attributes influence language models' internal representations of user expertise, which can affect model behavior despite no observable behavioral bias.
The introduction of a causal framework that decomposes occupational bias into internal representations and observable outputs.
The study may face challenges in reproducibility due to the complexity of causal frameworks and the variability in model architectures.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.