Papers/2609.30287
🧪 Test?View on arXiv

A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID

Not provided

neural networkstext detectionBERTAI ethics
2609.30287
Builder Relevance
70%
1h ago

Abstract

This study investigates the neurons in a frozen BERT model that are responsible for AI-text detection, revealing a small, stable set of neurons that retain high detection accuracy across different text generators.

Reality Card

Core Claim

The study identifies a stable set of under 1% of neurons in BERT that are crucial for AI-text detection, which can generalize across different text generators.

Method / Result

The selected neurons retain 86-94% of the full-feature ceiling on unseen generator families.

Limitations

The study's findings may be limited by the specific architecture of BERT and the reliance on a single benchmark (RAID).

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers