Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
Zixuan Huang, Yikun Ban, Lean Fu, Xiaojie Li, Zhongxiang Dai, Jianxin Li, Deqing Wang
Abstract
Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its performance is highly dependent on the quality of the underlying human preference data. To address this bottleneck, prior work has explored various data selection strategies, but these methods often overlook the impact of the evolving states of the language model during the optimization process. In this paper, we introduce a novel problem: Sample Scheduling for DPO, which aims to dynamically and adaptively schedule training samples based on the model's evolving batch-wise states throughout preference optimization. To solve this problem, we propose SamS, an efficient and effective algorithm that adaptively selects samples in each training batch based on the LLM's learning feedback to maximize the potential generalization performance. Notably, without modifying the core DPO algorithm, simply integrating SamS significantly improves performance across tasks, with minimal additional computational overhead. This work points to a promising new direction for improving LLM alignment through batch-wise sample selection, with potential generalization to RLHF and broader supervised learning paradigms. The code is available at https://github.com/hzx122/SamS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Prompt-SID: Learning Structural Representation Prompt via Latent Diffusion for Single Image DenoisingHuaqiu Li, Wang Zhang, Xiaowan Hu, Tao Jiang et al.AAAI 2025 · 8 citations
- Contextual Rollout Bandits for Reinforcement Learning with Verifiable RewardsXiaodong Lu, Xiaohan Wang, Jiajun Chai, Guojun Yin et al.ICML 2026 · 7 citations
- LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior SamplingHuaqiu Li, Yong Wang, Tongwen Huang, Hailang Huang et al.ICCV 2025 · 4 citations
- SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative RecommendationWei Chen, Xingyu Guo, Shuang Li, Fuwei Zhang et al.ICML 2026 · 2 citations
Builds on43
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
Related papers
- β-DPO: Direct Preference Optimization with Dynamic βJunkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu et al.NeurIPS 2024 · 114 citations
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen et al.NeurIPS 2025 · 13 citations
- Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference OptimizationJiahui Zhu, Yuanjie Shi, Xiyue Peng, Xin Liu et al.ICLR 2026
- Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay PerspectiveRuichen Shao, Bei Li, Gangao Liu, Yang Chen et al.ICLR 2025
- Expectation Preference Optimization: Reliable Preference Estimation for Improving the Reasoning Capability of Large Language ModelsZelin Li, Dawei SongEMNLP 2025
