Proximal Point Imitation Learning
Luca Viano, Angeliki Kamoutsi, Gergely Neu, Igor Krawczuk, Volkan Cevher
Abstract
This work develops new algorithms with rigorous efficiency guarantees for infinite horizon imitation learning (IL) with linear function approximation without restrictive coherence assumptions. We begin with the minimax formulation of the problem and then outline how to leverage classical tools from optimization, in particular, the proximal-point method (PPM) and dual smoothing, for online and offline IL, respectively. Thanks to PPM, we avoid nested policy evaluation and cost updates for online IL appearing in the prior literature. In particular, we do away with the conventional alternating updates by the optimization of a single convex and smooth objective over both cost and Q-functions. When solved inexactly, we relate the optimization errors to the suboptimality of the recovered policy. As an added bonus, by re-interpreting PPM as dual smoothing with the expert policy as a center point, we also obtain an offline IL algorithm enjoying theoretical guarantees in terms of required expert trajectories. Finally, we achieve convincing empirical performance for both linear and neural network function approximation. From the technical point of view, the most important related works are the analysis of REPS/Q-REPS [90, 14, 89] and O-REPS [124] that first pointed out the connection between REPS and PPM. We build on their techniques with some important differences. In particular, while in the LP formulation of RL, PPM and mirror descent [15, 47] are equivalent, recognizing that they are not equivalent in IL is critical for stronger empirical performance. As an independent interest, our techniques can be used to improve upon the best rate for REPS in the tabular setting [89] and to , where E π s denotes the expectation with respect to the trajectories generated by π starting from s 0 = s. The optimal value function , is known to characterize optimal behaviors. Indeed V ⋆ c is the unique solution to the Bellman optimality equation V ⋆ c (s) = min a Q ⋆ c (s, a). In addition, any deterministic policy π ⋆ c (s) = arg min a Q ⋆ c (s, a) is known to be optimal. For every policy π, we define the normalized state-action occupancy measure µ , where P π ν0 [•] denotes the probability of an event when following π starting from s 0 ∼ ν 0 . The occupancy measure can be interpreted as the discounted visitation frequency of state-action pairs. This allows us to write ρ c (π) = ⟨µ π , c⟩.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3efcd56c-464b-4821-920d-5266233f5e7bCited by top-tier papers14
- Dual RL: Unification and New Methods for Reinforcement and Imitation LearningHarshit Sikchi, Qinqing Zheng, Amy Zhang, Scott NiekumICLR 2024 · 48 citations
- Fast Imitation via Behavior Foundation ModelsMatteo Pirotta, Andrea Tirinzoni, Ahmed Touati, Alessandro Lazaric et al.ICLR 2024 · 26 citations
- Imitating Language via Scalable Inverse Reinforcement LearningMarkus Wulfmeier, Michael Bloesch, Nino Vieillard, Arun Ahuja et al.NeurIPS 2024 · 26 citations
- Coherent Soft Imitation LearningJoe Watson, Sandy H. Huang, Nicolas HeessNeurIPS 2023 · 26 citations
- Imitation Learning in Discounted Linear MDPs without exploration assumptionsLuca Viano, Stratis Skoulakis, Volkan CevherICML 2024 · 10 citations
Builds on30
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 304 citations
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPsAlekh Agarwal, Sham M. Kakade, Akshay Krishnamurthy, Wen SunNeurIPS 2020 · 271 citations
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
Related papers
- Imitation with Neural Density ModelsKuno Kim, Akshat Jindal, Yang Song, Jiaming Song et al.NeurIPS 2021 · 14 citations
- On the Value of Interaction and Function Approximation in Imitation LearningNived Rajaraman, Yanjun Han, Lin Yang, Jingbo Liu et al.NeurIPS 2021 · 28 citations
- Provably and Practically Efficient Adversarial Imitation Learning with General Function ApproximationTian Xu, Zhilong Zhang, Ruishuo Chen, Yihao Sun et al.NeurIPS 2024 · 8 citations
- Optimistic Planning by Regularized Dynamic ProgrammingAntoine Moulin, Gergely NeuICML 2023 · 8 citations
- A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPsKihyuk Hong, Ambuj TewariICML 2024 · 5 citations
