Leveraging Approximate Symbolic Models for Reinforcement Learning via Skill Diversity
Lin Guan, Sarath Sreedharan, Subbarao Kambhampati
摘要
Creating reinforcement learning (RL) agents that are capable of accepting and leveraging task-specific knowledge from humans has been long identified as a possible strategy for developing scalable approaches for solving long-horizon problems. While previous works have looked at the possibility of using symbolic models along with RL approaches, they tend to assume that the high-level action models are executable at low level and the fluents can exclusively characterize all desirable MDP states. Symbolic models of real world tasks are however often incomplete. To this end, we introduce Approximate Symbolic-Model Guided Reinforcement Learning, wherein we will formalize the relationship between the symbolic model and the underlying MDP that will allow us to characterize the incompleteness of the symbolic model. We will use these models to extract high-level landmarks that will be used to decompose the task. At the low level, we learn a set of diverse policies for each possible task subgoal identified by the landmark, which are then stitched together. We evaluate our system by testing on three different benchmark domains and show how even with incomplete symbolic model information, our approach is able to discover the task structure and efficiently guide the RL agent towards the goal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task PlanningLin Guan, Karthik Valmeekam, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 被引用 347 次
- AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion ModelZibin Dong, Yifu Yuan, Jianye Hao, Fei Ni 等ICLR 2024 · 被引用 44 次
- ACE: Off-Policy Actor-Critic with Causality-Aware Entropy RegularizationTianying Ji, Yongyuan Liang, Yan Zeng, Yu Luo 等ICML 2024 · 被引用 20 次
- Integrating Planning and Deep Reinforcement Learning via Automatic Induction of Task SubstructuresJung-Chun Liu, Chi-Hsien Chang, Shao-Hua Sun, Tian-Li YuICLR 2024 · 被引用 6 次
- Relative Behavioral Attributes: Filling the Gap between Symbolic Goal Specification and Reward Learning from Human PreferencesLin Guan, Karthik Valmeekam, Subbarao KambhampatiICLR 2023
它引用的顶会 Paper4
- Compositional Reinforcement Learning from Logical SpecificationsKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurNeurIPS 2021 · 被引用 112 次
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 被引用 109 次
- Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data AugmentationLin Guan, Mudit Verma, Sihang Guo, Ruohan Zhang 等NeurIPS 2021 · 被引用 57 次
- CMAX++ : Leveraging Experience in Planning and Execution using Inaccurate ModelsAnirudh Vemula, J. Andrew Bagnell, Maxim LikhachevAAAI 2021 · 被引用 10 次
相关 Paper
- Optimistic Exploration in Reinforcement Learning Using Symbolic Model EstimatesSarath Sreedharan, Michael KatzNeurIPS 2023 · 被引用 12 次
- A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward MachinesWeichao Zhou, Wenchao LiICML 2022 · 被引用 15 次
- Representing Partial Programs with Blended Abstract SemanticsMaxwell I. Nye, Yewen Pu, Matthew Bowers, Jacob Andreas 等ICLR 2021 · 被引用 23 次
- Landmark-Guided Subgoal Generation in Hierarchical Reinforcement LearningJunsu Kim, Younggyo Seo, Jinwoo ShinNeurIPS 2021 · 被引用 90 次
- Possibility Before Utility: Learning And Using Hierarchical AffordancesRobby Costales, Shariq Iqbal, Fei ShaICLR 2022 · 被引用 5 次
