Inverse Reinforcement Learning with the Average Reward Criterion
Feiyang Wu, Jingyang Ke, Anqi Wu
Abstract
We study the problem of Inverse Reinforcement Learning (IRL) with an average-reward criterion. The goal is to recover an unknown policy and a reward function when the agent only has samples of states and actions from an experienced agent. Previous IRL methods assume that the expert is trained in a discounted environment, and the discount factor is known. This work alleviates this assumption by proposing an average-reward framework with efficient learning algorithms. We develop novel stochastic first-order methods to solve the IRL problem under the average-reward setting, which requires solving an Average-reward Markov Decision Process (AMDP) as a subproblem. To solve the subproblem, we develop a Stochastic Policy Mirror Descent (SPMD) method under general state and action spaces that needs steps of gradient computation. Equipped with SPMD, we propose the Inverse Policy Mirror Descent (IPMD) method for solving the IRL problem with a complexity. To the best of our knowledge, the aforementioned complexity results are new in IRL. Finally, we corroborate our analysis with numerical experiments using the MuJoCo benchmark and additional control tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Imitation Learning in Discounted Linear MDPs without exploration assumptionsLuca Viano, Stratis Skoulakis, Volkan CevherICML 2024 · 10 citations
- Conformal Inverse OptimizationBo Lin, Erick Delage, Timothy C. Y. ChanNeurIPS 2024 · 7 citations
- On the Expressivity of Objective-Specification Formalisms in Reinforcement LearningRohan Subramani, Marcus Williams, Max Heitmann, Halfdan Holm et al.ICLR 2024 · 3 citations
- Inverse Reinforcement Learning with Switching Rewards and History Dependency for Characterizing Animal BehaviorsJingyang Ke, Feiyang Wu, Jiyi Wang, Jeffrey Markowitz et al.ICML 2025
- Provable Policy Gradient for Robust Average-Reward MDPs Beyond RectangularityQiuhao Wang, Yuqi Zha, Chin Pang Ho, Marek PetrikICML 2025
Builds on6
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Improving Sample Complexity Bounds for (Natural) Actor-Critic AlgorithmsTengyu Xu, Zhe Wang, Yingbin LiangNeurIPS 2020 · 110 citations
- Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time GuaranteesSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2022 · 60 citations
- Finite-Sample Analysis of Off-Policy Natural Actor-Critic AlgorithmSajad Khodadadian, Zaiwei Chen, Siva Theja MaguluriICML 2021 · 33 citations
- Towards Theoretical Understanding of Inverse Reinforcement LearningAlberto Maria Metelli, Filippo Lazzati, Marcello RestelliICML 2023 · 21 citations
Related papers
- Non-Adversarial Inverse Reinforcement Learning via Successor Feature MatchingArnav Kumar Jain, Harley Wiltzer, Jesse Farebrother, Irina Rish et al.ICLR 2025
- Is Inverse Reinforcement Learning Harder than Standard Reinforcement Learning? A Theoretical PerspectiveLei Zhao, Mengdi Wang, Yu BaiICML 2024 · 3 citations
- Distributional Inverse Reinforcement LearningFeiyang Wu, Ye Zhao, Anqi WuICML 2026 · 1 citation
- Active Exploration for Inverse Reinforcement LearningDavid Lindner, Andreas Krause, Giorgia RamponiNeurIPS 2022 · 36 citations
- Inverse Reinforcement Learning from a Gradient-based LearnerGiorgia Ramponi, Gianluca Drappo, Marcello RestelliNeurIPS 2020 · 16 citations
