Papers/2609.25021
🧪 Test?View on arXiv

"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It

Author1, Author2, Author3, Author4, Author5

self-referential voiceactivation steeringlanguage modelsAI safety
2609.25021
Builder Relevance
70%
1h ago

Abstract

This paper explores how chat templates influence the self-referential voice of large language models, revealing that the presence of these templates alters the models' disclaimers and experiential expressions.

Reality Card

Core Claim

The study demonstrates that chat templates act as a switch that modulates the self-referential voice of LLMs, affecting how they express disclaimers and experiences.

Method / Result

The research identifies a specific direction in the activation space that can steer the disclaimer voice, showing that removing or adding this direction alters the models' self-reports.

Limitations

The findings suggest that researchers must control for the influence of chat templates when studying LLM self-reports, indicating a potential confound in existing studies.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers