Token-level Direct Preference Optimization
Yongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang, Haifeng Zhang, Jun Wang
Abstract
Fine-tuning pre-trained Large Language Models (LLMs) is essential to align them with human values and intentions. This process often utilizes methods like pairwise comparisons and KL divergence against a reference LLM, focusing on the evaluation of full answers generated by the models. However, the generation of these responses occurs in a token level, following a sequential, auto-regressive fashion. In this paper, we introduce Token-level Direct Preference Optimization (TDPO), a novel approach to align LLMs with human preferences by optimizing policy at the token level. Unlike previous methods, which face challenges in divergence efficiency, TDPO incorporates forward KL divergence constraints for each token, improving alignment and diversity. Utilizing the Bradley-Terry model for a token-based reward system, TDPO enhances the regulation of KL divergence, while preserving simplicity without the need for explicit reward modeling. Experimental results across various text tasks demonstrate TDPO's superior performance in balancing alignment with generation diversity. Notably, fine-tuning with TDPO strikes a better balance than DPO in the controlled sentiment generation and single-turn dialogue datasets, and significantly improves the quality of generated responses compared to both DPO and PPO-based RLHF methods. Our code is open-sourced at https://github.com/Vance0124/Token-level-Direct-Preference-Optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a204e6dc-cb18-4be4-83e7-611d908cbe98Cited by top-tier papers72
- Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMsXuan Zhang, Chao Du, Tianyu Pang, Qian Liu et al.NeurIPS 2024 · 177 citations
- Fast Best-of-N Decoding via Speculative RejectionHanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang et al.NeurIPS 2024 · 144 citations
- Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization ApproachWeiyu Ma, Qirui Mi, Yongcheng Zeng, Xue Yan et al.NeurIPS 2024 · 122 citations
- Zebra-CoT: A Dataset for Interleaved Vision-Language ReasoningAng Li, Charles L. Wang, Deqing Fu, Kaiyu Yue et al.ICLR 2026 · 85 citations
- Cal-DPO: Calibrated Direct Preference Optimization for Language Model AlignmentTeng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li et al.NeurIPS 2024 · 76 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
Related papers
- Risk-aware Direct Preference Optimization under Nested Risk MeasureLijun Zhang, Lin Li, Yajie Qi, Huizhong Song et al.NeurIPS 2025 · 3 citations
- TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference OptimizationMingkang Zhu, Xi Chen, Zhongdao Wang, Bei Yu et al.ICML 2025
- TokenRatio: Principled Token-Level Preference Optimization via Ratio MatchingTruong Nguyen, Tien-Phat Nguyen, Linh Van, Duy Nguyen et al.ICML 2026
- TPO: Aligning Large Language Models with Multi-branch & Multi-step Preference TreesWeibin Liao, Xu Chu, Yasha WangICLR 2025
- AlignDistil: Token-Level Language Model Alignment as Adaptive Policy DistillationSongming Zhang, Xue Zhang, Tong Zhang, Bojie Hu et al.ACL 2025
