-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data
Jake Fawkes, Jason Hartford
Abstract
In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities is an effective, low variance, surrogate loss for training generative models. This loss has the property that when evaluated on-policy its gradients correspond to those of the KL divergence, while off-policy it remains a valid loss with the same global minimiser. In this work, we demonstrate that this construction can be extended to the whole family of -divergences, leading to a family of losses whose on-policy gradients are that of the corresponding -divergence, but retain the same global minimiser off-policy. Specifically, we show that the on-policy gradients lead to a one to one correspondence between translation invariant loss functions on the target and model log probabilities, and -divergences. This equivalence allows us to design new surrogate loss functions for tuning a wide class of generative models that inherit the properties of the corresponding -divergence, such as being more mode covering, whilst being applicable to off-policy data. We apply our losses on a range of tasks, including classic synthetic examples, SynFlowNets for molecule discovery, and asynchronous large language model (LLM) tuning, demonstrating that our models retain their predicted properties on- and off-policy and can be applied to a wide class of generative models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fc3a85e-efa0-476d-a28b-3ddce2659f95Builds on13
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup et al.NeurIPS 2021 · 565 citations
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun et al.NeurIPS 2022 · 316 citations
- Aligning Language Models with Preferences through f-divergence MinimizationDongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen et al.ICML 2023 · 119 citations
- A theory of continuous generative flow networksSalem Lahlou, Tristan Deleu, Pablo Lemos, Dinghuai Zhang et al.ICML 2023 · 118 citations
Related papers
- On Divergence Measures for Training GFlowNetsTiago da Silva, Eliezer de Souza da Silva, Diego MesquitaNeurIPS 2024 · 8 citations
- GFlowNets and variational inferenceNikolay Malkin, Salem Lahlou, Tristan Deleu, Xu Ji et al.ICLR 2023
- Loss Functions and Operators Generated by f-DivergencesVincent Roulet, Tianlin Liu, Nino Vieillard, Michael Eli Sander et al.ICML 2025
- Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow NetworksRui Hu, Yifan Zhang, Zhuoran Li, Longbo HuangICLR 2025
- f-Divergence Variational InferenceNeng Wan, Dapeng Li, Naira HovakimyanNeurIPS 2020 · 13 citations
