Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal
Not provided in the abstract
Abstract
The study investigates the stability of embedded steering in language models after fine-tuning, revealing that while the steering mechanism remains intact, its functional effectiveness can degrade.
Reality Card
Embedded steering mechanisms in language models are durable at the weight level but can lose functional effectiveness after fine-tuning, necessitating behavioral re-validation.
Refusal ablation loses 64% of its effect on average under supervised fine-tuning (SFT).
The study indicates that while the weight edit survives fine-tuning, the behavioral outcomes can significantly degrade, raising concerns about the reliability of embedded steering post-deployment.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.