Revisiting OOD Generalization in Programmatic RL
Amirhossein Rajabpour, Kiarash Aghakasiri, Sandra Zilles, Levi Lelis
Abstract
Programmatic policies are often reported to generalize better than neural policies in reinforcement learning (RL) benchmarks. We revisit some of these claims and show that much of the observed gap arises from uncontrolled experimental factors rather than intrinsic representational reasons. Re-evaluating three core benchmarks used in influential papers---TORCS, Karel, and Parking---we find that neural policies, when trained with a few modifications, such as sparse observations and cautious intrinsic reward functions, can match or exceed the out-of-distribution (OOD) generalization of programmatic policies. We argue that a representation enables OOD generalization if (i) the policy space it induces includes a generalizing policy and (ii) the search algorithm can find it. The neural and programmatic policies in prior work are comparable in OOD generalization because the domain-specific languages used induce policy spaces similar to those of neural networks, and our modifications help the gradient search find generalizing solutions. By disentangling representational factors from experimental confounds, we advance our understanding of what makes a representation succeed or fail at OOD generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d7a0f17-9f9f-4170-9966-8d92f5d56b34Builds on11
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 104 citations
- Look where you look! Saliency-guided Q-networks for generalization in visual Reinforcement LearningDavid Bertoin, Adil Zouitine, Mehdi Zouitine, Emmanuel RachelsonNeurIPS 2022 · 67 citations
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 54 citations
- Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic ObjectivesWenjie Qiu, Wensen Mao, He ZhuNeurIPS 2023 · 44 citations
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 42 citations
Related papers
- Addressing Optimism Bias in Sequence Modeling for Reinforcement LearningAdam R. Villaflor, Zhe Huang, Swapnil Pande, John M. Dolan et al.ICML 2022 · 30 citations
- Hierarchical Programmatic Reinforcement Learning via Learning to Compose ProgramsGuan-Ting Liu, En-Pei Hu, Pu-Jen Cheng, Hung-Yi Lee et al.ICML 2023 · 21 citations
- Bayesian Reparameterization of Reward-Conditioned Reinforcement Learning with Energy-based ModelsWenhao Ding, Tong Che, Ding Zhao, Marco PavoneICML 2023 · 3 citations
- QORA: Zero-Shot Transfer via Interpretable Object-Relational Model LearningGabriel Stella, Dmitri LoguinovICML 2024 · 1 citation
- When Data Geometry Meets Deep Function: Generalizing Offline Reinforcement LearningJianxiong Li, Xianyuan Zhan, Haoran Xu, Xiangyu Zhu et al.ICLR 2023 · 4 citations
