Generalization to New Actions in Reinforcement Learning
Ayush Jain, Andrew Szot, Joseph J. Lim
摘要
A fundamental trait of intelligence is the ability to achieve goals in the face of novel circumstances, such as making decisions from new action choices. However, standard reinforcement learning assumes a fixed set of actions and requires expensive retraining when given a new action set. To make learning agents more adaptable, we introduce the problem of zero-shot generalization to new actions. We propose a two-stage framework where the agent first infers action representations from action information acquired separately from the task. A policy flexible to varying action sets is then trained with generalization objectives. We benchmark generalization on sequential tasks, such as selecting from an unseen tool-set to solve physical reasoning puzzles and stacking towers with novel 3D shapes. Videos and code are available at https://sites.google.com/ view/action-generalization .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 被引用 104 次
- RODE: Learning Roles to Decompose Multi-Agent TasksTonghan Wang, Tarun Gupta, Anuj Mahajan, Bei Peng 等ICLR 2021 · 被引用 60 次
- Grounding Multimodal Large Language Models in ActionsAndrew Szot, Bogdan Mazoure, Harsh Agrawal, R. Devon Hjelm 等NeurIPS 2024 · 被引用 43 次
- In-Context Reinforcement Learning for Variable Action SpacesViacheslav Sinii, Alexander Nikulin, Vladislav Kurenkov, Ilya Zisman 等ICML 2024 · 被引用 26 次
- I-PHYRE: Interactive Physical ReasoningShiqian Li, Kewen Wu, Chi Zhang, Yixin ZhuICLR 2024 · 被引用 16 次
它引用的顶会 Paper2
相关 Paper
- AnyMorph: Learning Transferable Polices By Inferring Agent MorphologyBrandon Trabucco, Mariano Phielipp, Glen BersethICML 2022 · 被引用 37 次
- Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementZhi Wang, Li Zhang, Wenhao Wu, Yuanheng Zhu 等NeurIPS 2024 · 被引用 31 次
- Blocks Assemble! Learning to Assemble with Large-Scale Structured Reinforcement LearningSeyed Kamyar Seyed Ghasemipour, Satoshi Kataoka, Byron David, Daniel Freeman 等ICML 2022 · 被引用 36 次
- Unsupervised Zero-Shot Reinforcement Learning via Functional Reward EncodingsKevin Frans, Seohong Park, Pieter Abbeel, Sergey LevineICML 2024 · 被引用 26 次
- Explore to Generalize in Zero-Shot RLEv Zisselman, Itai Lavie, Daniel Soudry, Aviv TamarNeurIPS 2023 · 被引用 26 次
