Papers/2609.11956
🧪 Test?View on arXiv

Performance, Efficiency and Collapse -- Advantages and Challenges in Offline Post-training of Code LLMs

Not provided

fine-tuningreinforcement learningcode generation
2609.11956
Builder Relevance
80%
1h ago

Abstract

This paper explores the feasibility of offline post-training for code-generating LLMs using existing datasets to enhance performance without the need for new sample generation.

Reality Card

Core Claim

Offline reinforcement learning can significantly improve zero-shot code generation performance of LLMs with minimal training time and without online sampling.

Method / Result

Performance gains observed across models ranging from 0.5B to 7B parameters.

Limitations

The extent of improvement varies among model families, which may affect generalizability.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers