🧪 Test?View on arXiv
Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention
Author1, Author2, Author3, Author4, Author5
fine-tuningconfidence calibrationabstentionlanguage models
2608.26121
Builder Relevance
1h ago80%
Abstract
The paper explores the use of a model's internal confidence as a substitute for labelled datasets to improve its ability to abstain from answering when uncertain.
Reality Card
Core Claim
The study demonstrates that a model's own confidence can effectively replace labelled datasets for training abstention in factual question answering.
Method / Result
The label-free method matches the performance of label-supervised abstention-tuning across six open-weights models.
Limitations
The method struggles with confidently wrong facts, which it cannot identify as uncertain.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.