What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus
Not specified in the provided content
Abstract
This paper audits the ISOT/Kaggle 'Fake and Real News' corpus, revealing that high accuracy scores are misleading and primarily reflect source and topic separability rather than true veracity.
Reality Card
The study concludes that within-corpus scores quantify source and topic separability rather than veracity, and recommends using metadata-only, small-sample, and topic-disjoint baselines for future diagnostics.
Removing leakage channels only lowers F1 by 1.21 points (from 0.9935 to 0.9814).
The benchmark is partly degenerate due to the presence of disjoint subjects in the metadata, leading to misleading high accuracy scores.
Paper to code
Verified implementation resources so builders can test the paperβs claims instead of stopping at the abstract.