The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models
Usama A. Khan, John Doe, Jane Smith, Alice Johnson, Bob Brown
Abstract
This paper investigates how system prompts affect the computation in language models, revealing that different types of instructions lead to varying degrees of restructuring in model representations.
Reality Card
The study demonstrates that while system prompts are recognized at every layer of the model, their impact on computation is limited and varies significantly by instruction type.
The mean CKA correlation between restrictive and permissive safety instructions is 0.997, indicating they engage similar computational pathways.
The findings may not be fully reproducible across all model architectures and sizes, particularly beyond the tested range of 1.5B to 72B parameters.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.