Compositional Foundation Models for Hierarchical Planning
Anurag Ajay, Seungwook Han, Yilun Du, Shuang Li, Abhi Gupta, Tommi S. Jaakkola, Joshua B. Tenenbaum, Leslie Pack Kaelbling, Akash Srivastava, Pulkit Agrawal
Abstract
To make effective decisions in novel environments with long-horizon goals, it is crucial to engage in hierarchical reasoning across spatial and temporal scales. This entails planning abstract subgoal sequences, visually reasoning about the underlying plans, and executing actions in accordance with the devised plan through visual-motor control. We propose Compositional Foundation Models for Hierarchical Planning (HiP), a foundation model which leverages multiple expert foundation model trained on language, vision and action data individually jointly together to solve long-horizon tasks. We use a large language model to construct symbolic plans that are grounded in the environment through a large video diffusion model. Generated video plans are then grounded to visual-motor control, through an inverse dynamics model that infers actions from generated videos. To enable effective reasoning within this hierarchy, we enforce consistency between the models via iterative refinement. We illustrate the efficacy and adaptability of our approach in three different long-horizon table-top manipulation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed83972c-778c-4e93-a997-88b44b2c5458Cited by top-tier papers15
- Provable Guarantees for Generative Behavior Cloning: Bridging Low-Level Stability and High-Level BehaviorAdam Block, Ali Jadbabaie, Daniel Pfrommer, Max Simchowitz et al.NeurIPS 2023 · 44 citations
- Generative Hierarchical Materials SearchSherry Yang, Simon L. Batzner, Ruiqi Gao, Muratahan Aykol et al.NeurIPS 2024 · 23 citations
- DecisionNCE: Embodied Multimodal Representations via Implicit Preference LearningJianxiong Li, Jinliang Zheng, Yinan Zheng, Liyuan Mao et al.ICML 2024 · 19 citations
- Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level CompositionJiahang Cao, Yize Huang, Hanzhong Guo, Qiang Zhang et al.ICLR 2026 · 14 citations
- Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional VideosKumaranage Ravindu Yasas Nagasinghe, Honglu Zhou, Malitha Gunawardhana, Martin Renqiang Min et al.CVPR 2024 · 5 citations
Builds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Video Language PlanningYilun Du, Sherry Yang, Pete Florence, Fei Xia et al.ICLR 2024 · 161 citations
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal GenerationSuraj Nair, Chelsea FinnICLR 2020 · 152 citations
- Extendable Planning via Multiscale DiffusionChang Chen, Hany Hamed, Doojin Baek, Taegu Kang et al.AAAI 2026 · 3 citations
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsYucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen et al.ICML 2025
- FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World ModelChongkai Gao, Haozhuo Zhang, Zhixuan Xu, Zhehao Cai et al.ICLR 2025
