Papers/2609.25006
πŸ“– Read?View on arXiv

What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus

Not specified in the provided content

shortcut learningfake news detectionmodel evaluation
2609.25006
Builder Relevance
80%
1h ago

Abstract

This paper audits the ISOT/Kaggle 'Fake and Real News' corpus, revealing that high accuracy scores are misleading and primarily reflect source and topic separability rather than true veracity.

Reality Card

Core Claim

The study concludes that within-corpus scores quantify source and topic separability rather than veracity, and recommends using metadata-only, small-sample, and topic-disjoint baselines for future diagnostics.

Method / Result

Removing leakage channels only lowers F1 by 1.21 points (from 0.9935 to 0.9814).

Limitations

The benchmark is partly degenerate due to the presence of disjoint subjects in the metadata, leading to misleading high accuracy scores.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers