Papers/2608.26222
🧪 Test?View on arXiv

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

Author1, Author2, Author3, Author4, Author5

safety evaluationfuzzinglarge language modelsneural networks
2608.26222
Builder Relevance
80%
1h ago

Abstract

NeuronFuzz introduces a novel fuzzing framework that utilizes internal safety neurons for efficient safety evaluation of LLMs.

Reality Card

Core Claim

NeuronFuzz achieves a jailbreak discovery rate of 76-100% across five white-box source models, significantly outperforming existing methods.

Method / Result

Achieves a jailbreak discovery rate of 76-100%, outperforming baselines by up to 48 percentage points.

Limitations

The reliance on specific internal safety neurons may limit generalizability across different LLM architectures.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers