Can Agents Run Relay Race with Strangers? Generalization of RL to Out-of-Distribution Trajectories
Li-Cheng Lan, Huan Zhang, Cho-Jui Hsieh
Abstract
In this paper, we define, evaluate, and improve the relay-generalization'' performance of reinforcement learning (RL) agents on the out-of-distribution controllable'' states. Ideally, an RL agent that generally masters a task should reach its goal starting from any controllable state of the environment instead of memorizing a small set of trajectories. For example, a self-driving system should be able to take over the control from humans in the middle of driving and must continue to drive the car safely. To practically evaluate this type of generalization, we start the test agent from the middle of other independently well-trained stranger agents' trajectories. With extensive experimental evaluation, we show the prevalence of generalization failure on controllable states from stranger agents. For example, in the Humanoid environment, we observed that a well-trained Proximal Policy Optimization (PPO) agent, with only 3.9% failure rate during regular testing, failed on 81.6% of the states generated by well-trained stranger PPO agents. To improve "relay generalization," we propose a novel method called Self-Trajectory Augmentation (STA), which will reset the environment to the agent's old states according to the Q function during training. After applying STA to the Soft Actor Critic's (SAC) training procedure, we reduced the failure rate of SAC under relay-evaluation by more than three times in most settings without impacting agent performance and increasing the needed number of environment interactions. Our code is available at https://github.com/lan-lc/STA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cba248e9-8045-4a32-9110-65cf5cd4eb3eCited by top-tier papers2
- FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement LearningYuwei Fu, Haichao Zhang, Di Wu, Wei Xu et al.ICML 2024 · 31 citations
- LAGEA: Language Guided Embodied Agents for Robotic ManipulationAbdul Monaf Chowdhury, Akm Moshiur Rahman Mazumder, Safaeid Arib, Rabeya AkterICML 2026 · 2 citations
Builds on6
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen et al.NeurIPS 2020 · 362 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 212 citations
- The Difficulty of Passive Learning in Deep Reinforcement LearningGeorg Ostrovski, Pablo Samuel Castro, Will DabneyNeurIPS 2021 · 73 citations
- Adaptable Agent Populations via a Generative Model of PoliciesKenneth Derek, Phillip IsolaNeurIPS 2021 · 18 citations
Related papers
- Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid ControlWeidong Huang, Zhehan Li, Hangxin Liu, Biao Hou et al.ICLR 2026 · 4 citations
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
- Proximal Supervised Fine-TuningWenhong Zhu, Ruobing Xie, Rui Wang, Xingwu Sun et al.ICLR 2026 · 13 citations
- Cross-Trajectory Representation Learning for Zero-Shot Generalization in RLBogdan Mazoure, Ahmed M. Ahmed, R. Devon Hjelm, Andrey Kolobov et al.ICLR 2022 · 30 citations
- On Trajectory Augmentations for Off-Policy EvaluationGe Gao, Qitong Gao, Xi Yang, Song Ju et al.ICLR 2024 · 5 citations
