Improving LLM-Based Recommenders with Conservative Generative Flow Networks
Xuan Yu, Feng Niu, Rui Zhu, Yudong Zhang, Xu Wang, Yang Wang
摘要
Generative Flow Networks (GFlowNets) have recently been used to improve diversity and mitigate popularity bias in LLM-based recommender systems, yet most objectives are developed under online-style assumptions. In offline LLM-based recommendation, learning is constrained to a fixed logged dataset, yielding partial support over token transitions on the dataset-induced token-prefix DAG. Naively applying Sub-Trajectory Balance (SubTB) becomes non-identifiable and can arbitrarily allocate probability mass to unsupported regions. We formalize this failure and identify three sources of non-identifiability that induce distributional shift between the dataset-implied policy and the learned policy: (i) flow overestimation, (ii) forward mass leakage, and (iii) backward compensation. To address it, we propose CFlower, which introduces a conservative SubTB objective that explicitly penalizes unsupported forward flow mass, and combines it with dataset-constrained policy learning with on-policy sampling on the dataset-induced DAG for efficient training under offline constraints. Experiments on three Amazon recommendation datasets show that CFlower improves distributional matching and delivers a stronger accuracy--exposure trade-off than prior GFlowNet and SFT baselines, while serving as a more reliable reference policy for downstream RL fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup 等NeurIPS 2021 · 被引用 565 次
- Learning GFlowNets From Partial Episodes For Improved Convergence And StabilityKanika Madan, Jarrid Rector-Brooks, Maksym Korablyov, Emmanuel Bengio 等ICML 2023 · 被引用 138 次
相关 Paper
- Process-Supervised LLM Recommenders via Flow-guided TuningChongming Gao, Mengyao Gao, Chenxiao Fan, Shuai Yuan 等SIGIR 2025 · 被引用 6 次
- Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet TrainingXi Wang, Wenbo Lu, Shenji WanICML 2026 · 被引用 1 次
- COFlowNet: Conservative Constraints on Flows Enable High-Quality Candidate GenerationYudong Zhang, Xuan Yu, Xu Wang, Zhaoyang Sun 等ICLR 2025
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun 等NeurIPS 2022 · 被引用 316 次
- Towards Understanding and Improving GFlowNet TrainingMax W. Shen, Emmanuel Bengio, Ehsan Hajiramezanali, Andreas Loukas 等ICML 2023 · 被引用 81 次
