Papers/2608.26119
🧪 Test?View on arXiv

DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs

Art Kanke, Author 2, Author 3, Author 4, Author 5

rhetorical fallacieslanguage modelsbenchmarkingmodel evaluation
2608.26119
Builder Relevance
70%
1h ago

Abstract

This paper introduces DeflectBench, a benchmark for evaluating the ability of large language models to generate rhetorical fallacies on demand.

Reality Card

Core Claim

DeflectBench demonstrates that the refusal of models to generate rhetorical fallacies is primarily influenced by the structure of the request rather than the content of the claims.

Method / Result

A single prompt frame change can swing model refusal rates by nearly 100 percentage points.

Limitations

The behavior of models often results in labeled compliance rather than clean compliance, complicating the interpretation of results.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers