🧪 Test?View on arXiv
Probe Generalization as Subspace Selection for OOD Deception Detection
Author1, Author2, Author3, Author4, Author5
deception detectionout-of-distributionlinear probesprincipal components
2609.02893
Builder Relevance
1h ago70%
Abstract
The study demonstrates that projecting inputs onto a small subset of principal components from the training distribution enhances the out-of-distribution generalization of linear probes for deception detection.
Reality Card
Core Claim
By selecting transferable principal components, the study achieves a 78% improvement on Insider Trading Report and a 25% improvement on Sandbagging in out-of-distribution deception detection.
Method / Result
Closed the baseline-to-oracle gap by 78% on Insider Trading Report.
Limitations
The reliance on specific principal components may limit generalizability to other datasets or tasks.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.