Process-Supervised LLM Recommenders via Flow-guided Tuning
Chongming Gao, Mengyao Gao, Chenxiao Fan, Shuai Yuan, Wentao Shi, Xiangnan He
摘要
While large language models (LLMs) are increasingly adapted for recommendation systems via supervised fine-tuning (SFT), this approach amplifies popularity bias due to its likelihood maximization objective, compromising recommendation diversity and fairness. To address this, we present Flow-guided fine-tuning recommender (Flower), which replaces SFT with a Generative Flow Network (GFlowNet) [6] framework that enacts process supervision through token-level reward propagation. Flower's key innovation lies in decomposing item-level rewards into constituent token rewards, enabling direct alignment between token generation probabilities and their reward signals. This mechanism achieves three critical advancements: (1) popularity bias mitigation and fairness enhancement through empirical distribution matching, (2) preservation of diversity through GFlowNet's proportional sampling, and (3) flexible integration of personalized preferences via adaptable token rewards. Experiments demonstrate Flower's superior distribution-fitting capability and its significant advantages over traditional SFT in terms of accuracy, fairness, and diversity, highlighting its potential to improve LLM-based recommendation systems. The implementation is available via https://github.com/Mr- Peach0301/Flower.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Fine-grained List-wise Alignment for Generative Medication RecommendationChenxiao Fan, Chongming Gao, Wentao Shi, Yaxin Gong 等NeurIPS 2025 · 被引用 14 次
- Agentic Feedback Loop Modeling Improves Recommendation and User SimulationShihao Cai, Jizhi Zhang, Keqin Bao, Chongming Gao 等SIGIR 2025 · 被引用 13 次
- Offline RL by Reward-Weighted Fine-Tuning for Conversation OptimizationSubhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi 等NeurIPS 2025 · 被引用 12 次
- IGD: Token Decisiveness Modeling via Information Gain in LLMs for Personalized RecommendationZijie Lin, Yang Zhang, Xiaoyan Zhao, Fengbin Zhu 等NeurIPS 2025 · 被引用 12 次
- Uncertainty-aware Generative RecommendationChenxiao Fan, Chongming Gao, Yaxin Gong, Haoyan Liu 等KDD 2026 · 被引用 2 次
它引用的顶会 Paper12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchDan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue 等NeurIPS 2024 · 被引用 527 次
- Understanding the Effects of RLHF on LLM Generalisation and DiversityRobert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina 等ICLR 2024 · 被引用 332 次
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun 等NeurIPS 2022 · 被引用 316 次
相关 Paper
- Improving LLM-Based Recommenders with Conservative Generative Flow NetworksXuan Yu, Feng Niu, Rui Zhu, Yudong Zhang 等ICML 2026
- GFlowGR: Fine-tuning Generative Recommendation Frameworks with Generative Flow NetworksYejing Wang, Shengyu Zhou, Jinyu Lu, Qidong Liu 等SIGIR 2026 · 被引用 1 次
- Does LLM Focus on the Right Words? Mitigating Context Bias in LLM-based RecommendersBohao Wang, Jiawei Chen, Feng Liu, Changwang Zhang 等WWW 2026 · 被引用 1 次
- LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language ModelsTiesunlong Shen, Rui Mao, Jin Wang, Heming Sun 等AAAI 2026 · 被引用 2 次
- Bridging Semantic Understanding and Popularity Bias with LLMsRenqiang Luo, Dong Zhang, Yupeng Gao, Wen Shi 等WWW 2026
