Solving Compositional Reinforcement Learning Problems via Task Reduction
Yunfei Li, Yilin Wu, Huazhe Xu, Xiaolong Wang, Yi Wu
Abstract
We propose a novel learning paradigm, Self-Imitation via Reduction (SIR), for solving compositional reinforcement learning problems. SIR is based on two core ideas: task reduction and self-imitation. Task reduction tackles a hard-to-solve task by actively reducing it to an easier task whose solution is known by the RL agent. Once the original hard task is successfully solved by task reduction, the agent naturally obtains a self-generated solution trajectory to imitate. By continuously collecting and imitating such demonstrations, the agent is able to progressively expand the solved subspace in the entire task space. Experiment results show that SIR can significantly accelerate and improve learning on a variety of challenging sparse-reward continuous-control problems with compositional structures. Code and videos are available at https://sites.google.com/ view/sir-compositional .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bad5dff5-9eee-42cd-9de1-abb9ec1f4650Cited by top-tier papers7
- Environment Generation for Zero-Shot Compositional Reinforcement LearningIzzeddin Gur, Natasha Jaques, Yingjie Miao, Jongwook Choi et al.NeurIPS 2021 · 50 citations
- Phasic Self-Imitative Reduction for Sparse-Reward Goal-Conditioned Reinforcement LearningYunfei Li, Tian Gao, Jiaqi Yang, Huazhe Xu et al.ICML 2022 · 25 citations
- E-MAPP: Efficient Multi-Agent Reinforcement Learning with Parallel Program GuidanceCan Chang, Ni Mu, Jiajun Wu, Ling Pan et al.NeurIPS 2022 · 14 citations
- Transitive RL: Value Learning via Divide and ConquerSeohong Park, Aditya Oberai, Pranav Atreya, Sergey LevineICLR 2026 · 12 citations
- MeMo: Meaningful, Modular Controllers via Noise InjectionMegan Tjandrasuwita, Jie Xu, Armando Solar-Lezama, Wojciech MatusikNeurIPS 2024 · 1 citation
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal GenerationSuraj Nair, Chelsea FinnICLR 2020 · 152 citations
Related papers
- Inverse Reinforcement Learning without Reinforcement LearningGokul Swamy, David Wu, Sanjiban Choudhury, Drew Bagnell et al.ICML 2023 · 49 citations
- Modular Lifelong Reinforcement Learning via Neural CompositionJorge A. Mendez, Harm van Seijen, Eric EatonICLR 2022 · 51 citations
- Self-Adaptive Imitation Learning: Learning Tasks with Delayed Rewards from Sub-optimal DemonstrationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouAAAI 2022 · 14 citations
- Robust Subtask Learning for Compositional GeneralizationKishor Jothimurugan, Steve Hsu, Osbert Bastani, Rajeev AlurICML 2023 · 7 citations
- Watch, Try, Learn: Meta-Learning from Demonstrations and RewardsAllan Zhou, Eric Jang, Daniel Kappler, Alexander Herzog et al.ICLR 2020 · 53 citations
