Component and Dimension Sparsity in Transformer Refusal Mechanisms
Wang Research Lab
Abstract
This paper investigates the structured mechanisms behind refusal steering in large language models, revealing the sparsity in components and dimensions that contribute to effective steering.
Reality Card
Refusal behaviors in transformers are not diffusely encoded but are instead assembled through identifiable sparse component mechanisms, allowing for effective steering with minimal intervention.
Effective steering retains 85-98% of the component-mechanism baseline using approximately 50% of residual stream dimensions.
The mechanistic basis of the interventions remains poorly understood despite the release of code and experimental results.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.