LobsDICE: Offline Learning from Observation via Stationary Distribution Correction Estimation
Geon-Hyeong Kim, Jongmin Lee, Youngsoo Jang, Hongseok Yang, Kee-Eung Kim
Abstract
We consider the problem of learning from observation (LfO), in which the agent aims to mimic the expert's behavior from the state-only demonstrations by experts. We additionally assume that the agent cannot interact with the environment but has access to the action-labeled transition data collected by some agents with unknown qualities. This offline setting for LfO is appealing in many real-world scenarios where the ground-truth expert actions are inaccessible and the arbitrary environment interactions are costly or risky. In this paper, we present LobsDICE, an offline LfO algorithm that learns to imitate the expert policy via optimization in the space of stationary distributions. Our algorithm solves a single convex minimization problem, which minimizes the divergence between the two statetransition distributions induced by the expert and the agent policy. Through an extensive set of offline LfO tasks, we show that LobsDICE outperforms strong baseline methods. * Equal contribution. † Work done while the authors were students at KAIST. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3456e361-18bf-4fa0-94e5-948b189641e8Cited by top-tier papers17
- Fast Imitation via Behavior Foundation ModelsMatteo Pirotta, Andrea Tirinzoni, Ahmed Touati, Alessandro Lazaric et al.ICLR 2024 · 26 citations
- Imitation Learning from Imperfection: Theoretical Justifications and AlgorithmsZiniu Li, Tian Xu, Zeyu Qin, Yang Yu et al.NeurIPS 2023 · 26 citations
- SEABO: A Simple Search-Based Method for Offline Imitation LearningJiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu et al.ICLR 2024 · 17 citations
- SPRINQL: Sub-optimal Demonstrations driven Offline Imitation LearningHuy Hoang, Tien Mai, Pradeep VarakanthamNeurIPS 2024 · 12 citations
- GO-DICE: Goal-Conditioned Option-Aware Offline Imitation Learning via Stationary Distribution Correction EstimationAbhinav Jain, Vaibhav V. UnhelkarAAAI 2024 · 11 citations
Builds on14
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 239 citations
- OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationJongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau et al.ICML 2021 · 137 citations
Related papers
- IOSTOM: Offline Imitation Learning from Observations via State Transition Occupancy MatchingQuang Anh Pham, Janaka Chathuranga Brahmanage, Tien Mai, Akshat KumarNeurIPS 2025 · 2 citations
- Offline Imitation from Observation via Primal Wasserstein State Occupancy MatchingKai Yan, Alexander G. Schwing, Yu-Xiong WangICML 2024 · 3 citations
- DemoDICE: Offline Imitation Learning with Supplementary Imperfect DemonstrationsGeon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon et al.ICLR 2022 · 111 citations
- Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution MatchingLantao Yu, Tianhe Yu, Jiaming Song, Willie Neiswanger et al.AAAI 2023 · 29 citations
- Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy MatchingYecheng Jason Ma, Andrew Shen, Dinesh Jayaraman, Osbert BastaniICML 2022 · 49 citations
