"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It
Author1, Author2, Author3, Author4, Author5
Abstract
This paper explores how chat templates influence the self-referential voice of large language models, revealing that the presence of these templates alters the models' disclaimers and experiential expressions.
Reality Card
The study demonstrates that chat templates act as a switch that modulates the self-referential voice of LLMs, affecting how they express disclaimers and experiences.
The research identifies a specific direction in the activation space that can steer the disclaimer voice, showing that removing or adding this direction alters the models' self-reports.
The findings suggest that researchers must control for the influence of chat templates when studying LLM self-reports, indicating a potential confound in existing studies.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.