SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection
Not provided in the abstract
Abstract
The paper introduces a benchmark that evaluates the ability of LLMs to reject factual errors across languages, revealing significant inconsistencies in performance.
Reality Card
The SWORD benchmark reveals that LLMs achieve higher accuracy on semantically plausible distortions than on nonsensical substitutions, indicating a reliance on distributional familiarity over factual verification.
Cross-lingual performance gaps reach up to 28 percentage points in (East) Asian languages when models are presented with distorted statements.
The paper does not provide specific details on the reproducibility of the distortion generation process or the models tested.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.