Can Agents Run Relay Race with Strangers? Generalization of RL to Out-of-Distribution Trajectories
Li-Cheng Lan, Huan Zhang, Cho-Jui Hsieh
摘要
In this paper, we define, evaluate, and improve the relay-generalization'' performance of reinforcement learning (RL) agents on the out-of-distribution controllable'' states. Ideally, an RL agent that generally masters a task should reach its goal starting from any controllable state of the environment instead of memorizing a small set of trajectories. For example, a self-driving system should be able to take over the control from humans in the middle of driving and must continue to drive the car safely. To practically evaluate this type of generalization, we start the test agent from the middle of other independently well-trained stranger agents' trajectories. With extensive experimental evaluation, we show the prevalence of generalization failure on controllable states from stranger agents. For example, in the Humanoid environment, we observed that a well-trained Proximal Policy Optimization (PPO) agent, with only 3.9% failure rate during regular testing, failed on 81.6% of the states generated by well-trained stranger PPO agents. To improve "relay generalization," we propose a novel method called Self-Trajectory Augmentation (STA), which will reset the environment to the agent's old states according to the Q function during training. After applying STA to the Soft Actor Critic's (SAC) training procedure, we reduced the failure rate of SAC under relay-evaluation by more than three times in most settings without impacting agent performance and increasing the needed number of environment interactions. Our code is available at https://github.com/lan-lc/STA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement LearningYuwei Fu, Haichao Zhang, Di Wu, Wei Xu 等ICML 2024 · 被引用 31 次
- LAGEA: Language Guided Embodied Agents for Robotic ManipulationAbdul Monaf Chowdhury, Akm Moshiur Rahman Mazumder, Safaeid Arib, Rabeya AkterICML 2026 · 被引用 2 次
它引用的顶会 Paper6
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 305 次
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 被引用 212 次
- The Difficulty of Passive Learning in Deep Reinforcement LearningGeorg Ostrovski, Pablo Samuel Castro, Will DabneyNeurIPS 2021 · 被引用 73 次
- Adaptable Agent Populations via a Generative Model of PoliciesKenneth Derek, Phillip IsolaNeurIPS 2021 · 被引用 18 次
相关 Paper
- Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid ControlWeidong Huang, Zhehan Li, Hangxin Liu, Biao Hou 等ICLR 2026 · 被引用 4 次
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 被引用 111 次
- Proximal Supervised Fine-TuningWenhong Zhu, Ruobing Xie, Rui Wang, Xingwu Sun 等ICLR 2026 · 被引用 13 次
- Cross-Trajectory Representation Learning for Zero-Shot Generalization in RLBogdan Mazoure, Ahmed M. Ahmed, R. Devon Hjelm, Andrey Kolobov 等ICLR 2022 · 被引用 30 次
- On Trajectory Augmentations for Off-Policy EvaluationGe Gao, Qitong Gao, Xi Yang, Song Ju 等ICLR 2024 · 被引用 5 次
