CO-PILOT: COllaborative Planning and reInforcement Learning On sub-Task curriculum
Shuang Ao, Tianyi Zhou, Guodong Long, Qinghua Lu, Liming Zhu, Jing Jiang
Abstract
Goal-conditioned reinforcement learning (RL) usually suffers from sparse reward and inefficient exploration in long-horizon tasks. Planning can find the shortest path to a distant goal that provides dense reward/guidance but is inaccurate without a precise environment model. We show that RL and planning can collaboratively learn from each other to overcome their own drawbacks. In "CO-PILOT", a learnable path-planner and an RL agent produce dense feedback to train each other on a curriculum of tree-structured sub-tasks. Firstly, the planner recursively decomposes a long-horizon task to a tree of sub-tasks in a top-down manner, whose layers construct coarse-to-fine sub-task sequences as plans to complete the original task. The planning policy is trained to minimize the RL agent's cost of completing the sequence in each layer from top to bottom layers, which gradually increases the sub-tasks and thus forms an easy-to-hard curriculum for the planner. Next, a bottom-up traversal of the tree trains the RL agent from easier sub-tasks with denser rewards on bottom layers to harder ones on top layers and collects its cost on each sub-task train the planner in the next episode. CO-PILOT repeats this mutual training for multiple episodes before switching to a new task, so the RL agent and planner are fully optimized to facilitate each other's training. We compare CO-PILOT with RL (SAC, HER, PPO), planning (RRT*, NEXT, SGT), and their combination (SoRB) on navigation and continuous control tasks. CO-PILOT significantly improves the success rate and sample efficiency. Our code is available at https://github.com/Shuang-AO/CO-PILOT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7b8aaf0-f347-446a-965d-ce81f293989eCited by top-tier papers6
- Domain Generalization via Balancing Training Difficulty and Model CapabilityXueying Jiang, Jiaxing Huang, Sheng Jin, Shijian LuICCV 2023 · 27 citations
- CSOT: Curriculum and Structure-Aware Optimal Transport for Learning with Noisy LabelsWanxing Chang, Ye Shi, Jingya WangNeurIPS 2023 · 24 citations
- Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldYijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao et al.CVPR 2024 · 23 citations
- Continual Task Allocation in Meta-Policy Network via Sparse PromptingYijun Yang, Tianyi Zhou, Jing Jiang, Guodong Long et al.ICML 2023 · 14 citations
- EAT-C: Environment-Adversarial sub-Task Curriculum for Efficient Reinforcement LearningShuang Ao, Tianyi Zhou, Jing Jiang, Guodong Long et al.ICML 2022 · 6 citations
Builds on6
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu et al.ICLR 2021 · 222 citations
- Long-Horizon Visual Planning with Goal-Conditioned Hierarchical PredictorsKarl Pertsch, Oleh Rybkin, Frederik Ebert, Shenghao Zhou et al.NeurIPS 2020 · 96 citations
- Learning to Plan in High Dimensions via Neural Exploration-Exploitation TreesBinghong Chen, Bo Dai, Qinjie Lin, Guo Ye et al.ICLR 2020 · 60 citations
- Sub-Goal Trees a Framework for Goal-Based Reinforcement LearningTom Jurgenson, Or Avner, Edward Groshev, Aviv TamarICML 2020 · 48 citations
Related papers
- Imitating Graph-Based Planning with Goal-Conditioned PoliciesJunsu Kim, Younggyo Seo, Sungsoo Ahn, Kyunghwan Son et al.ICLR 2023 · 2 citations
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 14 citations
- CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement LearningUtsav Singh, Vinay P. NamboodiriAAAI 2026 · 6 citations
- Flexible and Efficient Long-Range Planning Through Curious ExplorationAidan Curtis, Minjian Xin, Dilip Arumugam, Kevin T. Feigelis et al.ICML 2020 · 7 citations
- C-Planning: An Automatic Curriculum for Learning Goal-Reaching TasksTianjun Zhang, Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine et al.ICLR 2022 · 19 citations
