Scaling up and Stabilizing Differentiable Planning with Implicit Differentiation
Linfeng Zhao, Huazhe Xu, Lawson L. S. Wong
Abstract
Differentiable planning promises end-to-end differentiability and adaptivity. However, an issue prevents it from scaling up to larger-scale problems: they need to differentiate through forward iteration layers to compute gradients, which couples forward computation and backpropagation, and needs to balance forward planner performance and computational cost of the backward pass. To alleviate this issue, we propose to differentiate through the Bellman fixed-point equation to decouple forward and backward passes for Value Iteration Network and its variants, which enables constant backward cost (in planning horizon) and flexible forward budget and helps scale up to large tasks. We study the convergence stability, scalability, and efficiency of the proposed implicit version of VIN and its variants and demonstrate their superiorities on a range of planning tasks: 2D navigation, visual navigation, and 2-DOF manipulation in configuration space and workspace.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5b7cba2-ea49-4d07-8d82-2250b033c65aCited by top-tier papers4
- Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact DynamicsShenao Zhang, Wanxin Jin, Zhaoran WangICML 2023 · 13 citations
- Integrating Symmetry into Differentiable Planning with Steerable ConvolutionsLinfeng Zhao, Xupeng Zhu, Lingzhi Kong, Robin Walters et al.ICLR 2023 · 1 citation
- DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit DifferentiationShuyuan Wang, Philip D. Loewen, Michael G. Forbes, R. Bhushan Gopaluni et al.ICML 2025
- Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term PlanningYuhui Wang, Qingyuan Wu, Dylan R. Ashley, Francesco Faccio et al.ICML 2025
Builds on13
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 272 citations
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 177 citations
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 129 citations
Related papers
- Towards real-world navigation with deep differentiable plannersShu Ishida, João F. HenriquesCVPR 2022 · 6 citations
- Highway Value Iteration NetworksYuhui Wang, Weida Li, Francesco Faccio, Qingyuan Wu et al.ICML 2024 · 3 citations
- Universal Value Iteration Networks: When Spatially-Invariant Is Not UniversalLi Zhang, Xin Li, Sen Chen, Hongyu Zang et al.AAAI 2020 · 5 citations
- : Implicit Layers for Implicit RepresentationsZhichun Huang, Shaojie Bai, J. Zico KolterNeurIPS 2021 · 5 citations
- Neural Algorithmic Reasoners are Implicit PlannersAndreea Deac, Petar Velickovic, Ognjen Milinkovic, Pierre-Luc Bacon et al.NeurIPS 2021 · 27 citations
