Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits
Author1, Author2, Author3, Author4, Author5
Abstract
This study explores how Large Language Models (LLMs) modulate the expression of Dark Triad traits under different social desirability conditions.
Reality Card
The study demonstrates that LLMs systematically adjust Dark Triad trait expressions based on contextual framing, with significant variations in response modulation across different models and traits.
Most models reduced Dark Triad scores under fake-good conditions and increased them under fake-bad conditions, with Machiavellianism and narcissism showing the strongest shifts.
The variability in response modulation across different models and traits may complicate reproducibility and generalization of findings.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.