Deep Inverse Q-learning with Constraints
Gabriel Kalweit, Maria Hügle, Moritz Werling, Joschka Boedecker
Abstract
Popular Maximum Entropy Inverse Reinforcement Learning approaches require the computation of expected state visitation frequencies for the optimal policy under an estimate of the reward function. This usually requires intermediate value estimation in the inner loop of the algorithm, slowing down convergence considerably. In this work, we introduce a novel class of algorithms that only needs to solve the MDP underlying the demonstrated behavior once to recover the expert policy. This is possible through a formulation that exploits a probabilistic behavior assumption for the demonstrations within the structure of Q-learning. We propose Inverse Actionvalue Iteration which is able to fully recover an underlying reward of an external agent in closed-form analytically. We further provide an accompanying class of sampling-based variants which do not depend on a model of the environment. We show how to extend this class of algorithms to continuous state-spaces via function approximation and how to estimate a corresponding action-value function, leading to a policy as close as possible to the policy of the external agent, while optionally satisfying a list of predefined hard constraints. We evaluate the resulting algorithms called Inverse Action-value Iteration, Inverse Q-learning and Deep Inverse Qlearning on the Objectworld benchmark, showing a speedup of up to several orders of magnitude compared to (Deep) Max-Entropy algorithms. We further apply Deep Constrained Inverse Q-learning on the task of learning autonomous lane-changes in the open-source simulator SUMO achieving competent driving after training on data corresponding to 30 minutes of demonstrations. * Equal Contribution. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 050d23c5-4261-49de-90e3-13442fb2805eCited by top-tier papers3
- Proximal Point Imitation LearningLuca Viano, Angeliki Kamoutsi, Gergely Neu, Igor Krawczuk et al.NeurIPS 2022 · 27 citations
- Learning Soft Constraints From Constrained Expert DemonstrationsAshish Gaurav, Kasra Rezaee, Guiliang Liu, Pascal PoupartICLR 2023 · 4 citations
- APC-RL: Exceeding data-driven behavior priors with adaptive policy compositionFinn Rietz, Pedro Zuidberg Dos Martires, Johannes A. StorkICLR 2026
Related papers
- Maximum Likelihood Constraint Inference for Inverse Reinforcement LearningDexter R. R. Scobee, S. Shankar SastryICLR 2020 · 74 citations
- Multi-Modal Inverse Constrained Reinforcement Learning from a Mixture of DemonstrationsGuanren Qiao, Guiliang Liu, Pascal Poupart, Zhiqiang XuNeurIPS 2023 · 28 citations
- Non-Adversarial Inverse Reinforcement Learning via Successor Feature MatchingArnav Kumar Jain, Harley Wiltzer, Jesse Farebrother, Irina Rish et al.ICLR 2025
- Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time GuaranteesSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2022 · 60 citations
- Distributed Inverse Constrained Reinforcement Learning for Multi-agent SystemsShicheng Liu, Minghui ZhuNeurIPS 2022 · 41 citations
