LS-IQ: Implicit Reward Regularization for Inverse Reinforcement Learning
Firas Al-Hafez, Davide Tateo, Oleg Arenz, Guoping Zhao, Jan Peters
Abstract
Recent methods for imitation learning directly learn a -function using an implicit reward formulation rather than an explicit reward function. However, these methods generally require implicit reward regularization to improve stability and often mistreat absorbing states. Previous works show that a squared norm regularization on the implicit reward function is effective, but do not provide a theoretical analysis of the resulting properties of the algorithms. In this work, we show that using this regularizer under a mixture distribution of the policy and the expert provides a particularly illuminating perspective: the original objective can be understood as squared Bellman error minimization, and the corresponding optimization problem minimizes a bounded -Divergence between the expert and the mixture distribution. This perspective allows us to address instabilities and properly treat absorbing states. We show that our method, Least Squares Inverse Q-Learning (LS-IQ), outperforms state-of-the-art algorithms, particularly in environments with absorbing states. Finally, we propose to use an inverse dynamics model to learn from observations only. Using this approach, we retain performance in settings where no expert actions are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 452c3152-040f-4b63-8f73-3f3726a808beCited by top-tier papers25
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 92 citations
- Dual RL: Unification and New Methods for Reinforcement and Imitation LearningHarshit Sikchi, Qinqing Zheng, Amy Zhang, Scott NiekumICLR 2024 · 48 citations
- Fast Imitation via Behavior Foundation ModelsMatteo Pirotta, Andrea Tirinzoni, Ahmed Touati, Alessandro Lazaric et al.ICLR 2024 · 26 citations
- Imitating Language via Scalable Inverse Reinforcement LearningMarkus Wulfmeier, Michael Bloesch, Nino Vieillard, Arun Ahuja et al.NeurIPS 2024 · 26 citations
- Coherent Soft Imitation LearningJoe Watson, Sandy H. Huang, Nicolas HeessNeurIPS 2023 · 26 citations
Builds on5
- Deep Evidential RegressionAlexander Amini, Wilko Schwarting, Ava Soleimany, Daniela RusNeurIPS 2020 · 777 citations
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine et al.SIGGRAPH 2021 · 392 citations
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 299 citations
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation GapGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Steven WuICML 2021 · 90 citations
Related papers
- Regularized Inverse Reinforcement LearningWonseok Jeon, Chen-Yang Su, Paul Barde, Thang Doan et al.ICLR 2021 · 6 citations
- Reward-free World Models for Online Imitation LearningShangzhe Li, Zhiao Huang, Hao SuICML 2025
- Robust Visual Imitation Learning with Inverse Dynamics RepresentationsSiyuan Li, Xun Wang, Rongchang Zuo, Kewu Sun et al.AAAI 2024 · 8 citations
- State-only Imitation with Transition Dynamics MismatchTanmay Gangwani, Jian PengICLR 2020 · 56 citations
- IOSTOM: Offline Imitation Learning from Observations via State Transition Occupancy MatchingQuang Anh Pham, Janaka Chathuranga Brahmanage, Tien Mai, Akshat KumarNeurIPS 2025 · 2 citations
