Distributional Inverse Reinforcement Learning
Feiyang Wu, Ye Zhao, Anqi Wu
摘要
We propose a distributional framework for offline Inverse Reinforcement Learning (IRL) that jointly models uncertainty over reward functions and full distributions of returns. Unlike conventional IRL approaches that recover a deterministic reward estimate or match only expected returns, our method captures richer structure in expert behavior, particularly in learning the reward distribution, by minimizing first-order stochastic dominance (FSD) violations and thus integrating distortion risk measures (DRMs) into policy learning, enabling the recovery of both reward distributions and distribution-aware policies. This formulation is well-suited for behavior analysis and risk-aware imitation learning. Theoretical analysis show that the algorithm converge with iteration complexity. Empirical results on synthetic benchmarks, real-world neurobehavioral data, and MuJoCo control tasks demonstrate that our method recovers expressive reward representations and achieves state-of-the-art imitation performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
- Risk-Averse Offline Reinforcement LearningNúria Armengol Urpí, Sebastian Curi, Andreas KrauseICLR 2021 · 被引用 81 次
- Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time GuaranteesSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2022 · 被引用 60 次
相关 Paper
- Inverse Reinforcement Learning with the Average Reward CriterionFeiyang Wu, Jingyang Ke, Anqi WuNeurIPS 2023 · 被引用 16 次
- Is Inverse Reinforcement Learning Harder than Standard Reinforcement Learning? A Theoretical PerspectiveLei Zhao, Mengdi Wang, Yu BaiICML 2024 · 被引用 3 次
- When Demonstrations meet Generative World Models: A Maximum Likelihood Framework for Offline Inverse Reinforcement LearningSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2023 · 被引用 33 次
- Recovering from Out-of-sample States via Inverse Dynamics in Offline Reinforcement LearningKe Jiang, Jia-Yu Yao, Xiaoyang TanNeurIPS 2023 · 被引用 12 次
- Learning Constraints from Offline Demonstrations via Superior Distribution Correction EstimationGuorui Quan, Zhiqiang Xu, Guiliang LiuICML 2024 · 被引用 4 次
