GFlowNet Training by Policy Gradients
Puhua Niu, Shili Wu, Mingzhou Fan, Xiaoning Qian
Abstract
Generative Flow Networks (GFlowNets) have been shown effective to generate combinatorial objects with desired properties. We here propose a new GFlowNet training framework, with policydependent rewards, that bridges keeping flow balance of GFlowNets to optimizing the expected accumulated reward in traditional Reinforcement-Learning (RL). This enables the derivation of new policy-based GFlowNet training methods, in contrast to existing ones resembling value-based RL. It is known that the design of backward policies in GFlowNet training affects efficiency. We further develop a coupled training strategy that jointly solves GFlowNet forward policy training and backward policy design. Performance analysis is provided with a theoretical guarantee of our policy-based GFlowNet training. Experiments on both simulated and real-world datasets verify that our policy-based strategies provide advanced RL perspectives for robust gradient estimation to improve GFlowNet performance. Our code is
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 012ee2be-f586-45fc-b3c0-b139c50e7c7dCited by top-tier papers4
- GFlowNet Assisted Biological Sequence EditingPouya M. Ghari, Alex M. Tseng, Gökcen Eraslan, Romain Lopez et al.NeurIPS 2024 · 10 citations
- Evaluating GFlowNet from partial episodes for stable and flexible policy-based trainingPuhua Niu, Shili Wu, Xiaoning QianICLR 2026 · 2 citations
- Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet TrainingXi Wang, Wenbo Lu, Shenji WanICML 2026 · 1 citation
- Random Policy Evaluation Uncovers Policies of Generative Flow NetworksHaoran He, Emmanuel Bengio, Qingpeng Cai, Ling PanICML 2025
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup et al.NeurIPS 2021 · 565 citations
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 343 citations
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun et al.NeurIPS 2022 · 316 citations
- Learning GFlowNets From Partial Episodes For Improved Convergence And StabilityKanika Madan, Jarrid Rector-Brooks, Maksym Korablyov, Emmanuel Bengio et al.ICML 2023 · 138 citations
Related papers
- Optimizing Backward Policies in GFlowNets via Trajectory Likelihood MaximizationTimofei Gritsaev, Nikita Morozov, Sergey Samsonov, Daniil TiapkinICLR 2025
- QGFN: Controllable Greediness with Action ValuesElaine Lau, Stephen Zhewen Lu, Ling Pan, Doina Precup et al.NeurIPS 2024 · 21 citations
- Towards Understanding and Improving GFlowNet TrainingMax W. Shen, Emmanuel Bengio, Ehsan Hajiramezanali, Andreas Loukas et al.ICML 2023 · 81 citations
- Pessimistic Backward Policy for GFlowNetsHyosoon Jang, Yunhui Jang, Minsu Kim, Jinkyoo Park et al.NeurIPS 2024 · 14 citations
- Flow Factorization for Efficient Generative Flow NetworksJiashun Liu, Chunhui Li, Cheng-Hao Liu, Dianbo Liu et al.AAAI 2025
