CEIP: Combining Explicit and Implicit Priors for Reinforcement Learning with Demonstrations
Kai Yan, Alexander G. Schwing, Yu-Xiong Wang
摘要
Although reinforcement learning has found widespread use in dense reward settings, training autonomous agents with sparse rewards remains challenging. To address this difficulty, prior work has shown promising results when using not only task-specific demonstrations but also task-agnostic albeit somewhat related demonstrations. In most cases, the available demonstrations are distilled into an implicit prior, commonly represented via a single deep net. Explicit priors in the form of a database that can be queried have also been shown to lead to encouraging results. To better benefit from available demonstrations, we develop a method to Combine Explicit and Implicit Priors (CEIP). CEIP exploits multiple implicit priors in the form of normalizing flows in parallel to form a single complex prior. Moreover, CEIP uses an effective explicit retrieval and push-forward mechanism to condition the implicit priors. In three challenging environments, we find the proposed CEIP method to improve upon sophisticated state-of-the-art techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- No Prior Mask: Eliminate Redundant Action for Deep Reinforcement LearningDianyu Zhong, Yiqin Yang, Qianchuan ZhaoAAAI 2024 · 被引用 15 次
- Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision TransformersKai Yan, Alexander G. Schwing, Yu-Xiong WangNeurIPS 2024 · 被引用 11 次
- How2Compress: Scalable and Efficient Edge Video Analytics via Adaptive Granular Video CompressionYuheng Wu, Thanh-Tung Nguyen, Lucas Liebe, Quang Tau 等ACM MM 2025 · 被引用 1 次
- Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline DataShilong Deng, Zetao Zheng, Hongcai He, Paul Weng 等AAAI 2025
它引用的顶会 Paper14
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin 等ICLR 2020 · 被引用 692 次
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin 等ICLR 2020 · 被引用 221 次
- Parrot: Data-Driven Behavioral Priors for Reinforcement LearningAvi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu 等ICLR 2021 · 被引用 161 次
- A Universal Representation Transformer Layer for Few-Shot Image ClassificationLu Liu, William L. Hamilton, Guodong Long, Jing Jiang 等ICLR 2021 · 被引用 143 次
- RotoGrad: Gradient Homogenization in Multitask LearningAdrián Javaloy, Isabel ValeraICLR 2022 · 被引用 114 次
相关 Paper
- APC-RL: Exceeding data-driven behavior priors with adaptive policy compositionFinn Rietz, Pedro Zuidberg Dos Martires, Johannes A. StorkICLR 2026
- CE-Nav: Flow-Guided Reinforcement Refinement for Cross-Embodiment Local NavigationKai Yang, Tianlin Zhang, Zhengbo Wang, Zedong Chu 等ICLR 2026 · 被引用 12 次
- Watch, Try, Learn: Meta-Learning from Demonstrations and RewardsAllan Zhou, Eric Jang, Daniel Kappler, Alexander Herzog 等ICLR 2020 · 被引用 53 次
- Preferential Normalizing FlowsPetrus Mikkola, Luigi Acerbi, Arto KlamiNeurIPS 2024 · 被引用 3 次
- Guided Exploration with Proximal Policy Optimization using a Single DemonstrationGabriele Libardi, Gianni De Fabritiis, Sebastian DittertICML 2021 · 被引用 32 次
