Papers/2608.16926
🧪 Test?View on arXiv

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

Not provided in the abstract

fine-tuningdata selectionmachine learningmodel optimization
2608.16926
Builder Relevance
80%
2h ago

Abstract

Data-DPO is a method for selecting effective samples from large-scale data to reduce training costs while maintaining model performance.

Reality Card

Core Claim

Data-DPO outperforms existing data selection baselines and achieves better performance than full data training under multiple data budgets.

Method / Result

Data-DPO consistently surpasses full data training performance across various data budgets.

Limitations

The abstract does not specify limitations or reproducibility concerns.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers