A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward Machines
Weichao Zhou, Wenchao Li
摘要
A misspecified reward can degrade sample efficiency and induce undesired behaviors in reinforcement learning (RL) problems. We propose symbolic reward machines for incorporating highlevel task knowledge when specifying the reward signals. Symbolic reward machines augment existing reward machine formalism by allowing transitions to carry predicates and symbolic reward outputs. This formalism lends itself well to inverse reinforcement learning, whereby the key challenge is determining appropriate assignments to the symbolic values from a few expert demonstrations. We propose a hierarchical Bayesian approach for inferring the most likely assignments such that the concretized reward machine can discriminate expert demonstrated trajectories from other trajectories with high accuracy. Experimental results show that learned reward machines can significantly improve training efficiency for complex RL tasks and generalize well across different task environment configurations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Modeling Others' Minds as CodeKunal Jha, Aydan Yuenan Huang, Eric Ye, Natasha Jaques 等ICLR 2026 · 被引用 6 次
- Rethinking Inverse Reinforcement Learning: from Data Alignment to Task AlignmentWeichao Zhou, Wenchao LiNeurIPS 2024 · 被引用 3 次
它引用的顶会 Paper6
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho 等NeurIPS 2021 · 被引用 107 次
- Adversarially Guided Actor-CriticYannis Flet-Berliac, Johan Ferret, Olivier Pietquin, Philippe Preux 等ICLR 2021 · 被引用 78 次
- Explicable Reward Design for Reinforcement Learning AgentsRati Devidze, Goran Radanovic, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2021 · 被引用 60 次
- Learning abstract structure for drawing by efficient motor program inductionLucas Yanan Tian, Kevin Ellis, Marta Kryven, Josh TenenbaumNeurIPS 2020 · 被引用 47 次
相关 Paper
- Programmatic Reward Design by ExampleWeichao Zhou, Wenchao LiAAAI 2022 · 被引用 15 次
- Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement LearningGuy Azran, Mohamad H. Danesh, Stefano V. Albrecht, Sarah KerenAAAI 2024 · 被引用 2 次
- Leveraging Approximate Symbolic Models for Reinforcement Learning via Skill DiversityLin Guan, Sarath Sreedharan, Subbarao KambhampatiICML 2022 · 被引用 31 次
- GRACE: A Language Model Framework for Explainable Inverse Reinforcement LearningSilvia Sapora, R. Devon Hjelm, Omar Attia, Alexander Toshev 等ICLR 2026
- Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited DataAndrew C. Li, Toryn Q. Klassen, Andrew Wang, Parand A. Alamdari 等NeurIPS 2025 · 被引用 5 次
