🧪 Test?View on arXiv
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
Art Kanke, Author 2, Author 3, Author 4, Author 5
rhetorical fallacieslanguage modelsbenchmarkingmodel evaluation
2608.26119
Builder Relevance
1h ago70%
Abstract
This paper introduces DeflectBench, a benchmark for evaluating the ability of large language models to generate rhetorical fallacies on demand.
Reality Card
Core Claim
DeflectBench demonstrates that the refusal of models to generate rhetorical fallacies is primarily influenced by the structure of the request rather than the content of the claims.
Method / Result
A single prompt frame change can swing model refusal rates by nearly 100 percentage points.
Limitations
The behavior of models often results in labeled compliance rather than clean compliance, complicating the interpretation of results.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.