Papers/2610.06903
🧪 Test?View on arXiv

Component and Dimension Sparsity in Transformer Refusal Mechanisms

Wang Research Lab

transformersrefusal mechanismssparsityactivation steering
2610.06903
Builder Relevance
70%
1h ago

Abstract

This paper investigates the structured mechanisms behind refusal steering in large language models, revealing the sparsity in components and dimensions that contribute to effective steering.

Reality Card

Core Claim

Refusal behaviors in transformers are not diffusely encoded but are instead assembled through identifiable sparse component mechanisms, allowing for effective steering with minimal intervention.

Method / Result

Effective steering retains 85-98% of the component-mechanism baseline using approximately 50% of residual stream dimensions.

Limitations

The mechanistic basis of the interventions remains poorly understood despite the release of code and experimental results.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers