LobsDICE: Offline Learning from Observation via Stationary Distribution Correction Estimation
Geon-Hyeong Kim, Jongmin Lee, Youngsoo Jang, Hongseok Yang, Kee-Eung Kim
摘要
We consider the problem of learning from observation (LfO), in which the agent aims to mimic the expert's behavior from the state-only demonstrations by experts. We additionally assume that the agent cannot interact with the environment but has access to the action-labeled transition data collected by some agents with unknown qualities. This offline setting for LfO is appealing in many real-world scenarios where the ground-truth expert actions are inaccessible and the arbitrary environment interactions are costly or risky. In this paper, we present LobsDICE, an offline LfO algorithm that learns to imitate the expert policy via optimization in the space of stationary distributions. Our algorithm solves a single convex minimization problem, which minimizes the divergence between the two statetransition distributions induced by the expert and the agent policy. Through an extensive set of offline LfO tasks, we show that LobsDICE outperforms strong baseline methods. * Equal contribution. † Work done while the authors were students at KAIST. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Fast Imitation via Behavior Foundation ModelsMatteo Pirotta, Andrea Tirinzoni, Ahmed Touati, Alessandro Lazaric 等ICLR 2024 · 被引用 26 次
- Imitation Learning from Imperfection: Theoretical Justifications and AlgorithmsZiniu Li, Tian Xu, Zeyu Qin, Yang Yu 等NeurIPS 2023 · 被引用 26 次
- SEABO: A Simple Search-Based Method for Offline Imitation LearningJiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu 等ICLR 2024 · 被引用 17 次
- SPRINQL: Sub-optimal Demonstrations driven Offline Imitation LearningHuy Hoang, Tien Mai, Pradeep VarakanthamNeurIPS 2024 · 被引用 12 次
- GO-DICE: Goal-Conditioned Option-Aware Offline Imitation Learning via Stationary Distribution Correction EstimationAbhinav Jain, Vaibhav V. UnhelkarAAAI 2024 · 被引用 11 次
它引用的顶会 Paper14
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
- OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationJongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau 等ICML 2021 · 被引用 137 次
相关 Paper
- IOSTOM: Offline Imitation Learning from Observations via State Transition Occupancy MatchingQuang Anh Pham, Janaka Chathuranga Brahmanage, Tien Mai, Akshat KumarNeurIPS 2025 · 被引用 2 次
- Offline Imitation from Observation via Primal Wasserstein State Occupancy MatchingKai Yan, Alexander G. Schwing, Yu-Xiong WangICML 2024 · 被引用 3 次
- DemoDICE: Offline Imitation Learning with Supplementary Imperfect DemonstrationsGeon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon 等ICLR 2022 · 被引用 111 次
- Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution MatchingLantao Yu, Tianhe Yu, Jiaming Song, Willie Neiswanger 等AAAI 2023 · 被引用 29 次
- Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy MatchingYecheng Jason Ma, Andrew Shen, Dinesh Jayaraman, Osbert BastaniICML 2022 · 被引用 49 次
