Papers/2609.02893
🧪 Test?View on arXiv

Probe Generalization as Subspace Selection for OOD Deception Detection

Author1, Author2, Author3, Author4, Author5

deception detectionout-of-distributionlinear probesprincipal components
2609.02893
Builder Relevance
70%
1h ago

Abstract

The study demonstrates that projecting inputs onto a small subset of principal components from the training distribution enhances the out-of-distribution generalization of linear probes for deception detection.

Reality Card

Core Claim

By selecting transferable principal components, the study achieves a 78% improvement on Insider Trading Report and a 25% improvement on Sandbagging in out-of-distribution deception detection.

Method / Result

Closed the baseline-to-oracle gap by 78% on Insider Trading Report.

Limitations

The reliance on specific principal components may limit generalizability to other datasets or tasks.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers