GFlowNet Training by Policy Gradients
Puhua Niu, Shili Wu, Mingzhou Fan, Xiaoning Qian
摘要
Generative Flow Networks (GFlowNets) have been shown effective to generate combinatorial objects with desired properties. We here propose a new GFlowNet training framework, with policydependent rewards, that bridges keeping flow balance of GFlowNets to optimizing the expected accumulated reward in traditional Reinforcement-Learning (RL). This enables the derivation of new policy-based GFlowNet training methods, in contrast to existing ones resembling value-based RL. It is known that the design of backward policies in GFlowNet training affects efficiency. We further develop a coupled training strategy that jointly solves GFlowNet forward policy training and backward policy design. Performance analysis is provided with a theoretical guarantee of our policy-based GFlowNet training. Experiments on both simulated and real-world datasets verify that our policy-based strategies provide advanced RL perspectives for robust gradient estimation to improve GFlowNet performance. Our code is
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- GFlowNet Assisted Biological Sequence EditingPouya M. Ghari, Alex M. Tseng, Gökcen Eraslan, Romain Lopez 等NeurIPS 2024 · 被引用 10 次
- Evaluating GFlowNet from partial episodes for stable and flexible policy-based trainingPuhua Niu, Shili Wu, Xiaoning QianICLR 2026 · 被引用 2 次
- Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet TrainingXi Wang, Wenbo Lu, Shenji WanICML 2026 · 被引用 1 次
- Random Policy Evaluation Uncovers Policies of Generative Flow NetworksHaoran He, Emmanuel Bengio, Qingpeng Cai, Ling PanICML 2025
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup 等NeurIPS 2021 · 被引用 565 次
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 被引用 343 次
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun 等NeurIPS 2022 · 被引用 316 次
- Learning GFlowNets From Partial Episodes For Improved Convergence And StabilityKanika Madan, Jarrid Rector-Brooks, Maksym Korablyov, Emmanuel Bengio 等ICML 2023 · 被引用 138 次
相关 Paper
- Optimizing Backward Policies in GFlowNets via Trajectory Likelihood MaximizationTimofei Gritsaev, Nikita Morozov, Sergey Samsonov, Daniil TiapkinICLR 2025
- QGFN: Controllable Greediness with Action ValuesElaine Lau, Stephen Zhewen Lu, Ling Pan, Doina Precup 等NeurIPS 2024 · 被引用 21 次
- Towards Understanding and Improving GFlowNet TrainingMax W. Shen, Emmanuel Bengio, Ehsan Hajiramezanali, Andreas Loukas 等ICML 2023 · 被引用 81 次
- Pessimistic Backward Policy for GFlowNetsHyosoon Jang, Yunhui Jang, Minsu Kim, Jinkyoo Park 等NeurIPS 2024 · 被引用 14 次
- Flow Factorization for Efficient Generative Flow NetworksJiashun Liu, Chunhui Li, Cheng-Hao Liu, Dianbo Liu 等AAAI 2025
