Papers/2609.15994
🧪 Test?View on arXiv

Latent Undertow: How Ordinary Typos Break Probes

Elad D. Koren, Author 2, Author 3, Author 4, Author 5

malicious prompt detectionlanguage modelsprobingtypo resilience
2609.15994
Builder Relevance
80%
1h ago

Abstract

This paper explores how ordinary typos affect the performance of probes designed to detect malicious prompts in language models.

Reality Card

Core Claim

Introducing a KV-cache fork significantly improves the detection capability of single-position probes against typos, closing 95% of the performance gap.

Method / Result

The proposed method reduces the residual performance gap to -0.6pp, which is an order of magnitude better than previous methods.

Limitations

The results may vary across different models, as the geometry of rotation-and-decay was replicated only on specific models (Llama-3.1-8B, Qwen3-8B, Gemma-4-E4B).

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers