GFlowGR: Fine-tuning Generative Recommendation Frameworks with Generative Flow Networks
Yejing Wang, Shengyu Zhou, Jinyu Lu, Qidong Liu, Xinhang Li, Wenlin Zhang, Feng Li, Pengjie Wang, Chuan Yu, Jian Xu, Bo Zheng, Xiangyu Zhao
摘要
Generative recommendation (GR) has shown great promise in industrial applications, particularly for candidate generation and end-to-end recommendations. However, existing GR training paradigms suffer from two fundamental mismatches with real-world deployment requirements. First, they optimize for point-wise prediction of a single ground-truth item, whereas practical systems must produce a diverse, high-value set of candidates. Second, they treat all user interactions as equally informative, ignoring their inherent differences in utility. Although reward-based fine-tuning offers a partial remedy, it often lacks token-level supervision. To address these challenges, we reformulate GR as a sequential set-generation problem and propose GFlowGR, a GFlowNet-based fine-tuning framework that explicitly aligns generation probabilities with item-level utilities. GFlowGR comprises three tightly integrated components, each addressing a key limitation of conventional fine-tuning: a trajectory sampler that constructs training trajectories from candidate sets to enable set-wise learning, a behavior-aware reward model that quantifies item utility to support value-aware optimization, and a GFlowNet objective that provides token-level supervision. Extensive experiments on three real-world datasets with two representative LLM-based GR backbones show consistent and significant improvements over strong baselines, validating the effectiveness of our approach. For real-world deployment, GFlowGR has been integrated into Taobao 's search advertising businesses, delivering a 0.4% relative improvement in annual revenue since its launch in mid-2025, corresponding to billion-level monetary gains. Code is available at https://github.com/Applied-Machine-Learning-Lab/SIGIR26_GFlowGR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- UNO! UNified Offline Training Paradigm for Learning Path RecommendationLinzhi Peng, Wentao Zhu, Ke Cheng, Heng Chang 等AAAI 2026
- Hierarchical Residual Policy Optimization for Generative RecommendationsKaifeng Guo, Yiming Yang, Jingtong Gao, Guolei Zeng 等KDD 2026
- Hyperbolic RQ-VAE enhanced Generative Recommendation with Differential-Length Codebook StrategyAoran Zhang, Yu-Bin Yang, Yonghong YuICML 2026
它引用的顶会 Paper25
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup 等NeurIPS 2021 · 被引用 565 次
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan 等NeurIPS 2023 · 被引用 474 次
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun 等NeurIPS 2022 · 被引用 316 次
- Biological Sequence Design with GFlowNetsMoksh Jain, Emmanuel Bengio, Alex Hernández-García, Jarrid Rector-Brooks 等ICML 2022 · 被引用 224 次
相关 Paper
- Process-Supervised LLM Recommenders via Flow-guided TuningChongming Gao, Mengyao Gao, Chenxiao Fan, Shuai Yuan 等SIGIR 2025 · 被引用 6 次
- LWGR: Lagrangian-Constrained Personalized World Knowledge for Generative RecommendationLingyu Mu, Hao Deng, Haibo Xing, Kaican Lin 等SIGIR 2026
- APAO: Bridging the Training-Inference Gap in Generative Recommendation via Adaptive Prefix-Aware OptimizationYuanqing Yu, Yifan Wang, Weizhi Ma, Zhiqiang Guo 等KDD 2026 · 被引用 4 次
- Generative Explore-Exploit: Training-free Optimization of Generative Recommender Systems using LLM OptimizersLütfi Kerem Senel, Besnik Fetahu, Davis Yoshida, Zhiyu Chen 等ACL 2024
- Does LLM Focus on the Right Words? Mitigating Context Bias in LLM-based RecommendersBohao Wang, Jiawei Chen, Feng Liu, Changwang Zhang 等WWW 2026 · 被引用 1 次
