Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required
Not specified in the provided content
Abstract
The paper argues that relying on a small set of coding benchmarks for evaluating coding capability leads to a misleading assessment of general coding ability.
Reality Card
Optimization for coding benchmarks does not translate to improved general coding capability, as evidenced by the lack of cross-task transfer and limited gains from benchmark optimization.
Post-trained checkpoints show little cross-task transfer, with no significant improvements on tasks or LiveCodeBench.
The reliance on a small number of benchmarks creates a gap in reliable evaluation standards, making reproducibility and generalization difficult.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.