🧪 Test?View on arXiv
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
Not specified in the provided content
safetymechanistic interpretabilityadversarial robustnessLLM architecture
2609.00051
Builder Relevance
1h ago80%
Abstract
This paper investigates the internal mechanisms of safety in Large Language Models (LLMs) and proposes a multi-stage safety circuit that enhances refusal behavior against harmful inputs.
Reality Card
Core Claim
Circuit-guided weight scaling improves safety rates under adversarial attacks by 26.5% with only a 1.7% drop in accuracy across standard benchmarks.
Method / Result
26.5% improvement in safety rates under attacks across six LLMs.
Limitations
The paper does not specify the authors or provide detailed experimental setups, which may hinder reproducibility.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.