Cross-Episodic Curriculum for Transformer Agents
Lucy Xiaoyang Shi, Yunfan Jiang, Jake Grigsby, Linxi Fan, Yuke Zhu
摘要
We present a new algorithm, Cross-Episodic Curriculum (CEC), to boost the learning efficiency and generalization of Transformer agents. Central to CEC is the placement of cross-episodic experiences into a Transformer's context, which forms the basis of a curriculum. By sequentially structuring online learning trials and mixed-quality demonstrations, CEC constructs curricula that encapsulate learning progression and proficiency increase across episodes. Such synergy combined with the potent pattern recognition capabilities of Transformer models delivers a powerful cross-episodic attention mechanism. The effectiveness of CEC is demonstrated under two representative scenarios: one involving multi-task reinforcement learning with discrete control, such as in DeepMind Lab, where the curriculum captures the learning progression in both individual and progressively complex settings, and the other involving imitation learning with mixed-quality data for continuous control, as seen in RoboMimic, where the curriculum captures the improvement in demonstrators' expertise. In all instances, policies resulting from CEC exhibit superior performance and strong generalization. Code is opensourced on the project website cec-agent.github.io to facilitate research on Transformer agent learning. 1 Following the canonical definition in Sutton and Barto [73] , we refer to the sequences of agent-environment interaction with clearly identified initial and terminal states as "episodes". We interchangeably use "episode", "trial", and "trajectory" in this work. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Reward Is Enough: LLMs Are In-Context Reinforcement LearnersKefan Song, Amir Moeini, Peng Wang, Lei Gong 等ICLR 2026 · 被引用 42 次
- Emergence of In-Context Reinforcement Learning from Noise DistillationIlya Zisman, Vladislav Kurenkov, Alexander Nikulin, Viacheslav Sinii 等ICML 2024 · 被引用 28 次
- AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with TransformersJake Grigsby, Justin Sasek, Samyak Parajuli, Daniel Adebi 等NeurIPS 2024 · 被引用 19 次
- Towards Provable Emergence of In-Context Reinforcement LearningJiuqi Wang, Rohan Chandra, Shangtong ZhangNeurIPS 2025 · 被引用 5 次
- AMAGO: Scalable In-Context Reinforcement Learning for Adaptive AgentsJake Grigsby, Linxi Fan, Yuke ZhuICLR 2024
它引用的顶会 Paper15
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- Behavior Transformers: Cloning modes with one stoneNur Muhammad Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, Lerrel PintoNeurIPS 2022 · 被引用 470 次
相关 Paper
- Efficient Cross-Episode Meta-RLGresa Shala, André Biedenkapp, Pierre Krack, Florian Walter 等ICLR 2025
- Transformers are Meta-Reinforcement LearnersLuckeciano C. MeloICML 2022 · 被引用 66 次
- SMART: Self-supervised Multi-task pretrAining with contRol TransformersYanchao Sun, Shuang Ma, Ratnesh Madaan, Rogerio Bonatti 等ICLR 2023 · 被引用 4 次
- TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement LearningGe Li, Dong Tian, Hongyi Zhou, Xinkai Jiang 等ICLR 2025
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 被引用 83 次
