Preference Alignment with Flow Matching
Minu Kim, Yongsik Lee, Sehyeok Kang, Jihwan Oh, Song Chong, Se-Young Yun
摘要
We present Preference Flow Matching (PFM), a new framework for preference-based reinforcement learning (PbRL) that streamlines the integration of preferences into an arbitrary class of pre-trained models. Existing PbRL methods require fine-tuning pre-trained models, which presents challenges such as scalability, inefficiency, and the need for model modifications, especially with black-box APIs like GPT-4. In contrast, PFM utilizes flow matching techniques to directly learn from preference data, thereby reducing the dependency on extensive fine-tuning of pre-trained models. By leveraging flow-based models, PFM transforms less preferred data into preferred outcomes, and effectively aligns model outputs with human preferences without relying on explicit or implicit reward function estimation, thus avoiding common issues like overfitting in reward models. We provide theoretical insights that support our method's alignment with standard PbRL objectives. Experimental results indicate the practical effectiveness of our method, offering a new direction in aligning a pre-trained model to preference. Our code is available at https://github.com/jadehaus/preference-flow-matching.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Mirror Flow Matching with Heavy-Tailed Priors for Generative Modeling on Convex DomainsYunrui Guan, Krishna Balasubramanian, Shiqian MaICLR 2026 · 被引用 10 次
- Steering Where to Diffuse: Generative Modeling of Phenotypic Response Simulation with Steered Diffusion BridgeRongchao Zhang, Chengxin Li, Yiwei Lou, Yuling Shi 等CVPR 2026 · 被引用 1 次
- FAVE: Flow-based Average Velocity Establishment for Sequential RecommendationKe Shi, Yao Zhang, Feng Guo, Jinyuan Zhang 等SIGIR 2026
- Handling Missing Responses under Cluster Dependence with Applications to Language Model EvaluationZhenghao Zeng, David Arbour, Avi Feller, Ishita Dasgupta 等NeurIPS 2025
它引用的顶会 Paper10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
相关 Paper
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka 等NeurIPS 2023 · 被引用 61 次
- PC-Flow: Preference Alignment in Flow Matching via ClassifierShaomeng Wang, He Wang, Longquan Dai, Jinhui TangAAAI 2026
- Would I Lie To You? Inference Time Alignment of Language Models using Direct Preference HeadsAvelina Asada Hadji-Kyriacou, Ognjen ArandjelovicNeurIPS 2024 · 被引用 6 次
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu 等AAAI 2024 · 被引用 357 次
- Value Gradient Guidance for Flow Matching AlignmentZhen Liu, Tim Z. Xiao, Carles Domingo-Enrich, Weiyang Liu 等NeurIPS 2025 · 被引用 15 次
