IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
Stefano Viel, Luca Viano, Volkan Cevher
Abstract
This paper introduces the SOAR framework for imitation learning. SOAR is an algorithmic template that learns a policy from expert demonstrations with a primal dual style algorithm that alternates cost and policy updates. Within the policy updates, the SOAR framework uses an actor critic method with multiple critics to estimate the critic uncertainty and build an optimistic critic fundamental to drive exploration. When instantiated in the tabular setting, we get a provable algorithm with guarantees that matches the best known results in the desired accuracy parameter ϵ. Practically, the SOAR template can boost the performance of any imitation learning algorithm based on Soft Actror Critic (SAC). As an example, we show that SOAR can boost consistently the performance of the following SAC-based imitation learning algorithms: f -IRL, ML-IRL and CSIL. Overall, thanks to SOAR, the required number of episodes to achieve the same performance is reduced by half. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on48
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 195 citations
- Epistemic Neural NetworksIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla et al.NeurIPS 2023 · 142 citations
- Toward the Fundamental Limits of Imitation LearningNived Rajaraman, Lin F. Yang, Jiantao Jiao, Kannan RamchandranNeurIPS 2020 · 137 citations
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 120 citations
Related papers
- Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double ExplorationHeyang Zhao, Xingrui Yu, David Mark Bossens, Ivor W. Tsang et al.ICLR 2025
- Policy Improvement via Imitation of Multiple OraclesChing-An Cheng, Andrey Kolobov, Alekh AgarwalNeurIPS 2020 · 36 citations
- SOAR: Supervision from Observation for Agentic Reinforcement LearningMeng Li, Lei Li, Xiting Wang, Yi Yuan et al.ACL 2026 · 1 citation
- Confidence-Aware Imitation Learning from Demonstrations with Varying OptimalitySongyuan Zhang, Zhangjie Cao, Dorsa Sadigh, Yanan SuiNeurIPS 2021 · 73 citations
- Active Policy Improvement from Multiple Black-box OraclesXuefeng Liu, Takuma Yoneda, Chaoqi Wang, Matthew R. Walter et al.ICML 2023 · 13 citations
