Learning Task Decomposition with Ordered Memory Policy Network
Yuchen Lu, Yikang Shen, Siyuan Zhou, Aaron C. Courville, Joshua B. Tenenbaum, Chuang Gan
Abstract
Many complex real-world tasks are composed of several levels of sub-tasks. Humans leverage these hierarchical structures to accelerate the learning process and achieve better generalization. In this work, we study the inductive bias and propose Ordered Memory Policy Network (OMPN) to discover subtask hierarchy by learning from demonstration. The discovered subtask hierarchy could be used to perform task decomposition, recovering the subtask boundaries in an unstructured demonstration. Experiments on Craft and Dial demonstrate that our model can achieve higher task decomposition performance under both unsupervised and weakly supervised settings, comparing with strong baselines. OMPN can also be directly applied to partially observable environments and still achieve higher task decomposition performance. Our visualization further confirms that the subtask hierarchy can emerge in our model 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc0254de-06ee-4081-9284-442edc6067dfCited by top-tier papers7
- Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal ReasoningWenhao Ding, Haohong Lin, Bo Li, Ding ZhaoNeurIPS 2022 · 59 citations
- Seeing is not Believing: Robust Reinforcement Learning against Spurious CorrelationWenhao Ding, Laixi Shi, Yuejie Chi, Ding ZhaoNeurIPS 2023 · 39 citations
- SCaR: Refining Skill Chaining for Long-Horizon Robotic Manipulation via Dual RegularizationZixuan Chen, Ze Ji, Jing Huo, Yang GaoNeurIPS 2024 · 26 citations
- Learning Options via CompressionYiding Jiang, Evan Zheran Liu, Benjamin Eysenbach, J. Zico Kolter et al.NeurIPS 2022 · 26 citations
- Possibility Before Utility: Learning And Using Hierarchical AffordancesRobby Costales, Shariq Iqbal, Fei ShaICLR 2022 · 5 citations
Builds on2
- Learning to Combine Top-Down and Bottom-Up Signals in Recurrent Neural Networks with Attention over ModulesSarthak Mittal, Alex Lamb, Anirudh Goyal, Vikram Voleti et al.ICML 2020 · 73 citations
- Learning Compound Tasks without Task-specific Knowledge via Imitation and Self-supervised LearningSang-Hyun Lee, Seung-Woo SeoICML 2020 · 12 citations
Related papers
- Chain-of-Thought Predictive ControlZhiwei Jia, Vineet Thumuluri, Fangchen Liu, Linghao Chen et al.ICML 2024 · 24 citations
- Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement LearningValerie Chen, Abhinav Gupta, Kenneth MarinoICLR 2021 · 6 citations
- Conditional Diffusion Model for Multi-Agent Dynamic Task DecompositionYanda Zhu, Yuanyang Zhu, Daoyi Dong, Caihua Chen et al.AAAI 2026
- Unsupervised Hierarchical Skill DiscoveryDamion Harvey, Geraud Nangue Tasse, Benjamin Rosman, Branden Ingram et al.ICML 2026 · 1 citation
- Skill Induction and Planning with Latent LanguagePratyusha Sharma, Antonio Torralba, Jacob AndreasACL 2022 · 127 citations
