🧪 Test?View on arXiv
Learning from the Gap Between Pass@K and Pass@1
Not specified in the provided content
fine-tuningreinforcement learninglanguage models
2609.35793
Builder Relevance
1h ago80%
Abstract
This paper introduces GapFT, a method that improves language model performance by fine-tuning on the gap between single-sample and multi-sample decoding outcomes.
Reality Card
Core Claim
GapFT improves Pass@1 accuracy by 14.4 and 13.9 points on LogiQA 2.0 and ReClor, respectively, while using only one third of the data compared to full verified fine-tuning.
Method / Result
GapFT matches the source model's verifier-selected Pass@4 accuracy with a single decode.
Limitations
The paper does not specify the authors or provide detailed reproducibility metrics for GapFT.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.
No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.