Preference Alignment with Flow Matching
Minu Kim, Yongsik Lee, Sehyeok Kang, Jihwan Oh, Song Chong, Se-Young Yun
Abstract
We present Preference Flow Matching (PFM), a new framework for preference-based reinforcement learning (PbRL) that streamlines the integration of preferences into an arbitrary class of pre-trained models. Existing PbRL methods require fine-tuning pre-trained models, which presents challenges such as scalability, inefficiency, and the need for model modifications, especially with black-box APIs like GPT-4. In contrast, PFM utilizes flow matching techniques to directly learn from preference data, thereby reducing the dependency on extensive fine-tuning of pre-trained models. By leveraging flow-based models, PFM transforms less preferred data into preferred outcomes, and effectively aligns model outputs with human preferences without relying on explicit or implicit reward function estimation, thus avoiding common issues like overfitting in reward models. We provide theoretical insights that support our method's alignment with standard PbRL objectives. Experimental results indicate the practical effectiveness of our method, offering a new direction in aligning a pre-trained model to preference. Our code is available at https://github.com/jadehaus/preference-flow-matching.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Mirror Flow Matching with Heavy-Tailed Priors for Generative Modeling on Convex DomainsYunrui Guan, Krishna Balasubramanian, Shiqian MaICLR 2026 · 10 citations
- Steering Where to Diffuse: Generative Modeling of Phenotypic Response Simulation with Steered Diffusion BridgeRongchao Zhang, Chengxin Li, Yiwei Lou, Yuling Shi et al.CVPR 2026 · 1 citation
- FAVE: Flow-based Average Velocity Establishment for Sequential RecommendationKe Shi, Yao Zhang, Feng Guo, Jinyuan Zhang et al.SIGIR 2026
- Handling Missing Responses under Cluster Dependence with Applications to Language Model EvaluationZhenghao Zeng, David Arbour, Avi Feller, Ishita Dasgupta et al.NeurIPS 2025
Builds on10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
Related papers
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka et al.NeurIPS 2023 · 61 citations
- PC-Flow: Preference Alignment in Flow Matching via ClassifierShaomeng Wang, He Wang, Longquan Dai, Jinhui TangAAAI 2026
- Would I Lie To You? Inference Time Alignment of Language Models using Direct Preference HeadsAvelina Asada Hadji-Kyriacou, Ognjen ArandjelovicNeurIPS 2024 · 6 citations
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu et al.AAAI 2024 · 357 citations
- Value Gradient Guidance for Flow Matching AlignmentZhen Liu, Tim Z. Xiao, Carles Domingo-Enrich, Weiyang Liu et al.NeurIPS 2025 · 15 citations
