Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization
Meng Li, Guangda Huzhang, Haibo Zhang, Xiting Wang, Anxiang Zeng
摘要
Direct Preference Optimization (DPO) has emerged as a promising framework for aligning Large Language Models (LLMs) with human preferences by directly optimizing the loglikelihood difference between chosen and rejected responses. However, existing methods assign equal importance to all tokens in the response, while humans focus on more meaningful parts. This leads to suboptimal preference optimization, as irrelevant or noisy tokens disproportionately influence DPO loss. To address this limitation, we propose Optimal Transportbased token weighting scheme for enhancing direct Preference Optimization (OTPO). By emphasizing semantically meaningful token pairs and de-emphasizing less relevant ones, our method introduces a context-aware token weighting scheme that yields a more contrastive reward difference estimate. This adaptive weighting enhances reward stability, improves interpretability, and ensures that preference optimization focuses on meaningful differences between responses. Extensive experiments have validated OTPO's effectiveness in improving instruction-following ability across various settings. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Revisiting Robustness for LLM Safety Alignment via Selective Geometry ControlYonghui Yang, Wenjian Tao, Jilong Liu, Xingyu Zhu 等ICML 2026 · 被引用 4 次
- AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly DetectionJunru Zhang, Lang Feng, Haoran Shi, Xu Guo 等ICML 2026
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
相关 Paper
- Token-Importance Guided Direct Preference OptimizationNing Yang, Hai Lin, Yibo Liu, Baoliang Tian 等ICLR 2026 · 被引用 14 次
- TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated WeightsAiwei Liu, Haoping Bai, Zhiyun Lu, Yanchao Sun 等ICLR 2025
- AlignDistil: Token-Level Language Model Alignment as Adaptive Policy DistillationSongming Zhang, Xue Zhang, Tong Zhang, Bojie Hu 等ACL 2025
- Autoregressive Direct Preference OptimizationMasanari Oi, Mahiro Ukai, Masahiro Kaneko, Naoaki Okazaki 等ICML 2026 · 被引用 1 次
- Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay PerspectiveRuichen Shao, Bei Li, Gangao Liu, Yang Chen 等ICLR 2025
