🧪 Test?View on arXiv
Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds
Not provided
token entanglementsubliminal learninglanguage modelscausal inference
2609.19149
Builder Relevance
2h ago70%
Abstract
This paper investigates how language models can convey hidden traits through seemingly unrelated outputs, focusing on the concept of token entanglement.
Reality Card
Core Claim
The study demonstrates that donor-control AUC significantly increases when transferring answer-position states between prompts, indicating a causal relationship in subliminal prompting.
Method / Result
Donor-control AUC rises from 0.254 to 0.540, a paired change of +0.286.
Limitations
The study's findings may not identify the exact mechanism of training-time trait transfer, limiting reproducibility.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.