Graph-GRPO: Training Graph Flow Models with Reinforcement Learning
Baoheng Zhu, Deyu Bo, Delvin Zhang, Xiao Wang
Abstract
Graph generation is a fundamental task with broad applications, such as drug discovery. Recently, discrete flow matching-based graph generation, a.k.a., graph flow model (GFM), has emerged due to its superior performance and flexible sampling. However, effectively aligning GFMs with complex human preferences or task-specific objectives remains a significant challenge. In this paper, we propose Graph-GRPO, an online reinforcement learning (RL) framework for training GFMs under verifiable rewards. Our method makes two key contributions: (1) We derive an analytical expression for the transition probability of GFMs, replacing the Monte Carlo sampling and enabling fully differentiable rollouts for RL training; (2) We propose a refinement strategy that randomly perturbs specific nodes and edges in a graph, and regenerates them, allowing for localized exploration and self-improvement of generation quality. Extensive experiments on both synthetic and real datasets demonstrate the effectiveness of Graph-GRPO. With only 50 denoising steps, our method achieves 95.0% and 97.5% Valid-Unique-Novelty scores on the planar and tree datasets, respectively. Moreover, Graph-GRPO achieves state-of-the-art performance on the molecular optimization tasks, outperforming graph-based and fragment-based RL methods as well as classic genetic algorithms. Code is available in https://github.com/Zhubaoheng/Graph-GRPO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8006d1a9-bfdf-49ad-97f8-469aaf93ce2aBuilds on26
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov et al.ICLR 2024 · 816 citations
- A Continuous Time Framework for Discrete Denoising ModelsAndrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth et al.NeurIPS 2022 · 496 citations
- Discrete Flow MatchingItai Gat, Tal Remez, Neta Shaul, Felix Kreuk et al.NeurIPS 2024 · 363 citations
- Score-based Generative Modeling of Graphs via the System of Stochastic Differential EquationsJaehyeong Jo, Seul Lee, Sung Ju HwangICML 2022 · 327 citations
Related papers
- GraphAF: a Flow-based Autoregressive Model for Molecular Graph GenerationChence Shi, Minkai Xu, Zhaocheng Zhu, Weinan Zhang et al.ICLR 2020 · 532 citations
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li et al.NeurIPS 2025 · 647 citations
- Fine-Grained GRPO for Precise Preference Alignment in Flow ModelsYujie Zhou, Pengyang Ling, Jiazi Bu, Yibin Wang et al.CVPR 2026 · 19 citations
- iGRPO: Fast Online RL for Flow Matching Model with Instant RewardSucheng Ren, Chen Chen, Zhenbang Wang, Liangchen Song et al.ICML 2026
- TEMPFLOW-GRPO: WHEN TIMING MATTERS FOR GRPO IN FLOW MODELSXiaoxuan He, Siming Fu, Yuke Zhao, Wanli Li et al.ICLR 2026 · 98 citations
