Papers/2608.26121
🧪 Test?View on arXiv

Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention

Author1, Author2, Author3, Author4, Author5

fine-tuningconfidence calibrationabstentionlanguage models
2608.26121
Builder Relevance
80%
1h ago

Abstract

The paper explores the use of a model's internal confidence as a substitute for labelled datasets to improve its ability to abstain from answering when uncertain.

Reality Card

Core Claim

The study demonstrates that a model's own confidence can effectively replace labelled datasets for training abstention in factual question answering.

Method / Result

The label-free method matches the performance of label-supervised abstention-tuning across six open-weights models.

Limitations

The method struggles with confidently wrong facts, which it cannot identify as uncertain.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers