-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data
Jake Fawkes, Jason Hartford
摘要
In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities is an effective, low variance, surrogate loss for training generative models. This loss has the property that when evaluated on-policy its gradients correspond to those of the KL divergence, while off-policy it remains a valid loss with the same global minimiser. In this work, we demonstrate that this construction can be extended to the whole family of -divergences, leading to a family of losses whose on-policy gradients are that of the corresponding -divergence, but retain the same global minimiser off-policy. Specifically, we show that the on-policy gradients lead to a one to one correspondence between translation invariant loss functions on the target and model log probabilities, and -divergences. This equivalence allows us to design new surrogate loss functions for tuning a wide class of generative models that inherit the properties of the corresponding -divergence, such as being more mode covering, whilst being applicable to off-policy data. We apply our losses on a range of tasks, including classic synthetic examples, SynFlowNets for molecule discovery, and asynchronous large language model (LLM) tuning, demonstrating that our models retain their predicted properties on- and off-policy and can be applied to a wide class of generative models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationEmmanuel Bengio, Moksh Jain, Maksym Korablyov, Doina Precup 等NeurIPS 2021 · 被引用 565 次
- Trajectory balance: Improved credit assignment in GFlowNetsNikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun 等NeurIPS 2022 · 被引用 316 次
- Aligning Language Models with Preferences through f-divergence MinimizationDongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen 等ICML 2023 · 被引用 119 次
- A theory of continuous generative flow networksSalem Lahlou, Tristan Deleu, Pablo Lemos, Dinghuai Zhang 等ICML 2023 · 被引用 118 次
相关 Paper
- On Divergence Measures for Training GFlowNetsTiago da Silva, Eliezer de Souza da Silva, Diego MesquitaNeurIPS 2024 · 被引用 8 次
- GFlowNets and variational inferenceNikolay Malkin, Salem Lahlou, Tristan Deleu, Xu Ji 等ICLR 2023
- Loss Functions and Operators Generated by f-DivergencesVincent Roulet, Tianlin Liu, Nino Vieillard, Michael Eli Sander 等ICML 2025
- Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow NetworksRui Hu, Yifan Zhang, Zhuoran Li, Longbo HuangICLR 2025
- f-Divergence Variational InferenceNeng Wan, Dapeng Li, Naira HovakimyanNeurIPS 2020 · 被引用 13 次
