EAT-C: Environment-Adversarial sub-Task Curriculum for Efficient Reinforcement Learning
Shuang Ao, Tianyi Zhou, Jing Jiang, Guodong Long, Xuan Song, Chengqi Zhang
Abstract
Reinforcement learning (RL) is inefficient on long-horizon tasks due to sparse rewards and its policy can be fragile to slightly perturbed environments. We address these challenges via a curriculum of tasks with coupled environments, generated by two policies trained jointly with RL: (1) a co-operative planning policy recursively decomposing a hard task into a coarse-to-fine sub-task tree; and (2) an adversarial policy modifying the environment in each sub-task. They are complementary to acquire more informative feedback for RL: (1) provides dense reward of easier subtasks while (2) modifies sub-tasks' environments to be more challenging and diverse. Conversely, they are trained by RL's dense feedback on subtasks so their generated curriculum keeps adaptive to RL's progress. The sub-task tree enables an easy-to-hard curriculum for every policy: its topdown construction gradually increases sub-tasks the planner needs to generate, while the adversarial training between the environment and RL follows a bottom-up traversal that starts from a dense sequence of easier sub-tasks allowing more frequent environment changes. We compare EAT-C with RL/planning targeting similar problems and methods with environment generators or adversarial agents. Extensive experiments on diverse tasks demonstrate the advantages of our method on improving RL's efficiency and generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e579530-843d-46d6-a332-0e9b3b8a6b38Cited by top-tier papers4
- Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldYijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao et al.CVPR 2024 · 23 citations
- Continual Task Allocation in Meta-Policy Network via Sparse PromptingYijun Yang, Tianyi Zhou, Jing Jiang, Guodong Long et al.ICML 2023 · 14 citations
- Robust Adversarial Reinforcement Learning via Bounded Rationality CurriculaAryaman Reddi, Maximilian Tölle, Jan Peters, Georgia Chalvatzaki et al.ICLR 2024 · 11 citations
- Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement LearningZeyang Liu, Lipeng Wan, Xinrui Yang, Zhuoran Chen et al.AAAI 2024 · 7 citations
Builds on7
- Automatic Curriculum Learning through Value DisagreementYunzhi Zhang, Pieter Abbeel, Lerrel PintoNeurIPS 2020 · 132 citations
- Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement LearningTianren Zhang, Shangqi Guo, Tian Tan, Xiaolin Hu et al.NeurIPS 2020 · 112 citations
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 97 citations
- Long-Horizon Visual Planning with Goal-Conditioned Hierarchical PredictorsKarl Pertsch, Oleh Rybkin, Frederik Ebert, Shenghao Zhou et al.NeurIPS 2020 · 96 citations
- Sub-Goal Trees a Framework for Goal-Based Reinforcement LearningTom Jurgenson, Or Avner, Edward Groshev, Aviv TamarICML 2020 · 48 citations
Related papers
- CO-PILOT: COllaborative Planning and reInforcement Learning On sub-Task curriculumShuang Ao, Tianyi Zhou, Guodong Long, Qinghua Lu et al.NeurIPS 2021 · 23 citations
- Understanding the Complexity Gains of Single-Task RL with a CurriculumQiyang Li, Yuexiang Zhai, Yi Ma, Sergey LevineICML 2023 · 21 citations
- CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement LearningUtsav Singh, Vinay P. NamboodiriAAAI 2026 · 6 citations
- Adaptive Procedural Task Generation for Hard-Exploration ProblemsKuan Fang, Yuke Zhu, Silvio Savarese, Li Fei-FeiICLR 2021 · 36 citations
- Beyond Fixed Tasks: Unsupervised Environment Design for Task-Level PairsDaniel Furelos-Blanco, Charles Pert, Frederik Kelbel, Alex F. Spies et al.AAAI 2026
