Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
Ruishuo Chen, Xun Wang, Rui Hu, Zhuoran Li, Longbo Huang
摘要
Generative Flow Networks (GFlowNets) excel at sampling diverse, high-reward objects. In many practical applications where active reward queries are infeasible, these models must be trained using static offline datasets. Prevailing training methods typically rely on a proxy model to provide reward feedback for online sampled trajectories. However, constructing a reliable proxy is often challenging due to data scarcity or high evaluation costs. While existing proxy-free approaches attempt to address this, they often impose coarse constraints that limit the model's ability to explore effectively. To overcome these limitations, we propose Trajectory-Distilled GFlowNet (TD-GFN), a novel proxy-free training framework. TD-GFN utilizes inverse reinforcement learning (IRL) to extract dense, transition-level edge rewards from offline trajectories, providing rich structural guidance for efficient exploration. Crucially, to ensure robustness, these rewards guide the policy indirectly through DAG pruning and prioritized backward sampling. This design ensures that gradient updates rely exclusively on ground-truth terminal rewards from the dataset, thereby preventing error propagation. Empirical results demonstrate that TD-GFN significantly outperforms a broad range of existing baselines in both convergence speed and sample quality, establishing a more robust and efficient paradigm for offline GFlowNet training. Recent studies, such as RO-GFlowNets (Wang et al., 2023) and COFlowNet (Zhang et al., 2025b), have explored learning directly from offline trajectories to eliminate dependence on proxy models. However, these approaches typically impose coarse constraints to align the policy with the dataset. Such practices can restrict generalization, inhibit effective
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper35
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun 等NeurIPS 2022 · 被引用 316 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
相关 Paper
- COFlowNet: Conservative Constraints on Flows Enable High-Quality Candidate GenerationYudong Zhang, Xuan Yu, Xu Wang, Zhaoyang Sun 等ICLR 2025
- Looking Backward: Retrospective Backward Synthesis for Goal-Conditioned GFlowNetsHaoran He, Can Chang, Huazhe Xu, Ling PanICLR 2025
- Pre-Training and Fine-Tuning Generative Flow NetworksLing Pan, Moksh Jain, Kanika Madan, Yoshua BengioICLR 2024 · 被引用 24 次
- GFlowNet Training by Policy GradientsPuhua Niu, Shili Wu, Mingzhou Fan, Xiaoning QianICML 2024 · 被引用 6 次
- Avoid What You Know: Divergent Trajectory Balance for GFlowNetsPedro Dall’Antonia, Tiago Silva, Daniel Csillag, Salem Lahlou 等ICML 2026 · 被引用 2 次
