Regularized Behavior Cloning for Blocking the Leakage of Past Action Information
Seokin Seo, HyeongJoo Hwang, Hongseok Yang, Kee-Eung Kim
摘要
For partially observable environments, imitation learning with observation histories (ILOH) assumes that control-relevant information is sufficiently captured in the observation histories for imitating the expert actions. In the offline setting where the agent is required to learn to imitate without interaction with the environment, behavior cloning (BC) has been shown to be a simple yet effective method for imitation learning. However, when the information about the actions executed in the past timesteps leaks into the observation histories, ILOH via BC often ends up imitating its own past actions. In this paper, we address this catastrophic failure by proposing a principled regularization for BC, which we name Past Action Leakage Regularization (PALR). The main idea behind our approach is to leverage the classical notion of conditional independence to mitigate the leakage. We compare different instances of our framework with natural choices of conditional independence metric and its estimator. The result of our comparison advocates the use of a particular kernel-based estimator for the conditional independence metric. We conduct an extensive set of experiments on benchmark datasets in order to assess the effectiveness of our regularization method. The experimental results show that our method significantly outperforms prior related approaches, highlighting its potential to successfully imitate expert actions when the past action information leaks into the observation histories.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- HAMLET: Switch Your Vision-Language-Action Model into a History-Aware PolicyMyungkyu Koo, Daewon Choi, Taeyoung Kim, Kyungmin Lee 等ICLR 2026 · 被引用 52 次
- Causal Action Influence Aware Counterfactual Data AugmentationNúria Armengol Urpí, Marco Bagatella, Marin Vlastelica, Georg MartiusICML 2024 · 被引用 11 次
- To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable RLYuda Song, Dhruv Rohatgi, Aarti Singh, J. Andrew BagnellNeurIPS 2025
它引用的顶会 Paper11
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 被引用 666 次
- A Measure-Theoretic Approach to Kernel Conditional Mean EmbeddingsJunhyung Park, Krikamol MuandetNeurIPS 2020 · 被引用 123 次
- DemoDICE: Offline Imitation Learning with Supplementary Imperfect DemonstrationsGeon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon 等ICLR 2022 · 被引用 111 次
- Discriminator-Weighted Offline Imitation Learning from Suboptimal DemonstrationsHaoran Xu, Xianyuan Zhan, Honglei Yin, Huiling QinICML 2022 · 被引用 105 次
相关 Paper
- Curriculum Offline Imitating LearningMinghuan Liu, Hanye Zhao, Zhengyu Yang, Jian Shen 等NeurIPS 2021 · 被引用 5 次
- Object-Aware Regularization for Addressing Causal Confusion in Imitation LearningJongjin Park, Younggyo Seo, Chang Liu, Li Zhao 等NeurIPS 2021 · 被引用 31 次
- A Simple Solution for Offline Imitation from Observations and Examples with Possibly Incomplete TrajectoriesKai Yan, Alexander G. Schwing, Yu-Xiong WangNeurIPS 2023 · 被引用 7 次
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 被引用 112 次
- Mutual Information Regularized Offline Reinforcement LearningXiao Ma, Bingyi Kang, Zhongwen Xu, Min Lin 等NeurIPS 2023 · 被引用 14 次
