Selective Preference Optimization via Token-Level Reward Function Estimation
Kailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang, Erxue Min, Sophia Ananiadou
摘要
Recent advancements in LLM alignment leverage token-level supervisions to perform finegrained preference optimization. However, existing token-level alignment methods either optimize on all available tokens, which can be noisy and inefficient, or perform selective training with complex and expensive key token selection strategies. In this work, we propose Selective Preference Optimization (SePO), a novel selective alignment strategy that centers on efficient key token selection without requiring strong, fine-grained supervision signals. We prove the feasibility of Direct Preference Optimization (DPO) as token-level reward function estimators, which applies to any existing alignment datasets and enables costefficient token selection with small-scale model sizes and training data. We then train an oracle model with DPO on the target data and utilize the estimated reward function to score all tokens within the target dataset, where only the key tokens are selected to supervise the target policy model with a contrastive objective function. Extensive experiments on three public evaluation benchmarks show that SePO significantly outperforms competitive baseline methods by only optimizing on 30% key tokens with up to 60% reduction in GPU training hours. We also explore SePO as a new paradigm for weakto-strong generalization, showing that weak oracle models effectively supervise strong policy models with up to 16.8× more parameters. SePO also selects useful supervision signals from out-of-distribution data, alleviating the over-optimization problem. The project is open-sourced here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language ModelsTiesunlong Shen, Rui Mao, Jin Wang, Heming Sun 等AAAI 2026 · 被引用 2 次
- Alignment-Aware DecodingFrédéric Berdoz, Luca Lanzendörfer, René Caky, Roger WattenhoferICML 2026 · 被引用 1 次
- VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video ModelsHaojian Huang, Haodong Chen, Shengqiong Wu, Meng Luo 等ICML 2025
- TLDR: Token-Level Detective Reward Model for Large Vision Language ModelsDeqing Fu, Tong Xiao, Rui Wang, Wang Zhu 等ICLR 2025
- ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference OptimizationHee Suk Yoon, Eunseop Yoon, Mark A. Hasegawa-Johnson, Sungwoong Kim 等ICML 2025
它引用的顶会 Paper19
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri 等NeurIPS 2023 · 被引用 516 次
相关 Paper
- AlignDistil: Token-Level Language Model Alignment as Adaptive Policy DistillationSongming Zhang, Xue Zhang, Tong Zhang, Bojie Hu 等ACL 2025
- Token-level Direct Preference OptimizationYongcheng Zeng, Guoqing Liu, Weiyu Ma, Ning Yang 等ICML 2024 · 被引用 136 次
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen 等NeurIPS 2025 · 被引用 13 次
- TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated WeightsAiwei Liu, Haoping Bai, Zhiyun Lu, Yanchao Sun 等ICLR 2025
- DPO Meets PPO: Reinforced Token Optimization for RLHFHan Zhong, Zikang Shan, Guhao Feng, Wei Xiong 等ICML 2025
