🧪 Test?View on arXiv
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
Author1, Author2, Author3, Author4, Author5
safety evaluationfuzzinglarge language modelsneural networks
2608.26222
Builder Relevance
1h ago80%
Abstract
NeuronFuzz introduces a novel fuzzing framework that utilizes internal safety neurons for efficient safety evaluation of LLMs.
Reality Card
Core Claim
NeuronFuzz achieves a jailbreak discovery rate of 76-100% across five white-box source models, significantly outperforming existing methods.
Method / Result
Achieves a jailbreak discovery rate of 76-100%, outperforming baselines by up to 48 percentage points.
Limitations
The reliance on specific internal safety neurons may limit generalizability across different LLM architectures.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.