Scaling up and Stabilizing Differentiable Planning with Implicit Differentiation
Linfeng Zhao, Huazhe Xu, Lawson L. S. Wong
摘要
Differentiable planning promises end-to-end differentiability and adaptivity. However, an issue prevents it from scaling up to larger-scale problems: they need to differentiate through forward iteration layers to compute gradients, which couples forward computation and backpropagation, and needs to balance forward planner performance and computational cost of the backward pass. To alleviate this issue, we propose to differentiate through the Bellman fixed-point equation to decouple forward and backward passes for Value Iteration Network and its variants, which enables constant backward cost (in planning horizon) and flexible forward budget and helps scale up to large tasks. We study the convergence stability, scalability, and efficiency of the proposed implicit version of VIN and its variants and demonstrate their superiorities on a range of planning tasks: 2D navigation, visual navigation, and 2-DOF manipulation in configuration space and workspace.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact DynamicsShenao Zhang, Wanxin Jin, Zhaoran WangICML 2023 · 被引用 13 次
- Integrating Symmetry into Differentiable Planning with Steerable ConvolutionsLinfeng Zhao, Xupeng Zhu, Lingzhi Kong, Robin Walters 等ICLR 2023 · 被引用 1 次
- DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit DifferentiationShuyuan Wang, Philip D. Loewen, Michael G. Forbes, R. Bhushan Gopaluni 等ICML 2025
- Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term PlanningYuhui Wang, Qingyuan Wu, Dylan R. Ashley, Francesco Faccio 等ICML 2025
它引用的顶会 Paper13
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 被引用 272 次
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 被引用 177 次
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 被引用 164 次
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 被引用 129 次
相关 Paper
- Towards real-world navigation with deep differentiable plannersShu Ishida, João F. HenriquesCVPR 2022 · 被引用 6 次
- Highway Value Iteration NetworksYuhui Wang, Weida Li, Francesco Faccio, Qingyuan Wu 等ICML 2024 · 被引用 3 次
- Universal Value Iteration Networks: When Spatially-Invariant Is Not UniversalLi Zhang, Xin Li, Sen Chen, Hongyu Zang 等AAAI 2020 · 被引用 5 次
- : Implicit Layers for Implicit RepresentationsZhichun Huang, Shaojie Bai, J. Zico KolterNeurIPS 2021 · 被引用 5 次
- Neural Algorithmic Reasoners are Implicit PlannersAndreea Deac, Petar Velickovic, Ognjen Milinkovic, Pierre-Luc Bacon 等NeurIPS 2021 · 被引用 27 次
