Temporal Difference Flows
Jesse Farebrother, Matteo Pirotta, Andrea Tirinzoni, Rémi Munos, Alessandro Lazaric, Ahmed Touati
摘要
Predictive models of the future are fundamental for an agent's ability to reason and plan. A common strategy learns a world model and unrolls it step-by-step at inference, where small errors can rapidly compound. Geometric Horizon Models (GHMs) offer a compelling alternative by directly making predictions of future states, avoiding cumulative inference errors. While GHMs can be conveniently learned by a generative analog to temporal difference (TD) learning, existing methods are negatively affected by bootstrapping predictions at train time and struggle to generate high-quality predictions at long horizons. This paper introduces Temporal Difference Flows (TD-Flow), which leverages the structure of a novel Bellman equation on probability paths alongside flow-matching techniques to learn accurate GHMs at over 5× the horizon length of prior methods. Theoretically, we establish a new convergence result and primarily attribute TD-Flow's efficacy to reduced gradient variance during training. We further show that similar arguments can be extended to diffusion-based methods. Empirically, we validate TD-Flow across a diverse set of domains on both generative metrics and downstream tasks, including policy evaluation. Moreover, integrating TD-Flow with recent behavior foundation models for planning over policies demonstrates substantial performance gains, underscoring its promise for long-horizon decision-making.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RLBhavya Agrawalla, Michal Nauman, Khush Agrawal, Aviral KumarICLR 2026 · 被引用 22 次
- Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-TuningRiccardo De Santi, Marin Vlastelica, Ya-Ping Hsieh, Zebang Shen 等NeurIPS 2025 · 被引用 17 次
- Value FlowsPerry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh 等ICLR 2026 · 被引用 13 次
- Intention-Conditioned Flow Occupancy ModelsChongyi Zheng, Seohong Park, Sergey Levine, Benjamin EysenbachICLR 2026 · 被引用 9 次
- Towards Robust Zero-Shot Reinforcement LearningKexin Zheng, Lauriane Teyssier, Yinan Zheng, Yu Luo 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 被引用 388 次
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
- Multisample Flow Matching: Straightening Flows with Minibatch CouplingsAram-Alexandre Pooladian, Heli Ben-Hamu, Carles Domingo-Enrich, Brandon Amos 等ICML 2023 · 被引用 243 次
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 被引用 206 次
相关 Paper
- HDFlow: Hierarchical Diffusion-Flow Planning for Long-horizon TasksGireesh Nandiraju, Yuanliang(Avery) Ju, Chaoyi Xu, Weiheng Liu 等ICML 2026
- Compositional Planning with Jumpy World ModelsJesse Farebrother, Matteo Pirotta, Andrea Tirinzoni, Marc Bellemare 等ICML 2026 · 被引用 1 次
- Gamma-Models: Generative Temporal Difference Learning for Infinite-Horizon PredictionMichael Janner, Igor Mordatch, Sergey LevineNeurIPS 2020 · 被引用 49 次
- Offline Reinforcement Learning with Universal Horizon ModelsHojun Chung, Junseo Lee, Songhwai OhICML 2026 · 被引用 1 次
- Generalised Policy Improvement with Geometric Policy CompositionShantanu Thakoor, Mark Rowland, Diana Borsa, Will Dabney 等ICML 2022 · 被引用 11 次
