Deep Laplacian-based Options for Temporally-Extended Exploration
Martin Klissarov, Marlos C. Machado
摘要
Selecting exploratory actions that generate a rich stream of experience for better learning is a fundamental challenge in reinforcement learning (RL). An approach to tackle this problem consists in selecting actions according to specific policies for an extended period of time, also known as options. A recent line of work to derive such exploratory options builds upon the eigenfunctions of the graph Laplacian. Importantly, until now these methods have been mostly limited to tabular domains where (1) the graph Laplacian matrix was either given or could be fully estimated, (2) performing eigendecomposition on this matrix was computationally tractable, and (3) value functions could be learned exactly. Additionally, these methods required a separate option discovery phase. These assumptions are fundamentally not scalable. In this paper we address these limitations and show how recent results for directly approximating the eigenfunctions of the Laplacian can be leveraged to truly scale up options-based exploration. To do so, we introduce a fully online deep RL algorithm for discovering Laplacian-based options and evaluate our approach on a variety of pixel-based tasks. We compare to several state-of-the-art exploration methods and show that our approach is effective, general, and especially promising in non-stationary settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu 等ICLR 2024 · 被引用 97 次
- METRA: Scalable Unsupervised RL with Metric-Aware AbstractionSeohong Park, Oleh Rybkin, Sergey LevineICLR 2024 · 被引用 83 次
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 被引用 72 次
- CQM: Curriculum Reinforcement Learning with a Quantized World ModelSeungjae Lee, Daesol Cho, Jonghae Park, H. Jin KimNeurIPS 2023 · 被引用 18 次
- SkiLD: Unsupervised Skill Discovery Guided by Factor InteractionsZizhao Wang, Jiaheng Hu, Caleb Chuck, Stephen Chen 等NeurIPS 2024 · 被引用 15 次
它引用的顶会 Paper13
- Option Discovery using Deep Skill ChainingAkhil Bagaria, George KonidarisICLR 2020 · 被引用 126 次
- Lipschitz-constrained Unsupervised Skill DiscoverySeohong Park, Jongwook Choi, Jaekyeom Kim, Honglak Lee 等ICLR 2022 · 被引用 72 次
- Exploration in Reinforcement Learning with Deep Covering OptionsYuu Jinnai, Jee Won Park, Marlos C. Machado, George Dimitri KonidarisICLR 2020 · 被引用 64 次
- Options of Interest: Temporal Abstraction with Interest FunctionsKhimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon 等AAAI 2020 · 被引用 51 次
- Unsupervised Skill Discovery with Bottleneck Option LearningJaekyeom Kim, Seohong Park, Gunhee KimICML 2021 · 被引用 39 次
相关 Paper
- Option Discovery in the Absence of Rewards with Manifold AnalysisAmitay Bar, Ronen Talmon, Ron MeirICML 2020 · 被引用 6 次
- Novel Exploration via OrthogonalityAndreas Theophilou, Özgür SimsekNeurIPS 2025 · 被引用 1 次
- Proper Laplacian Representation LearningDiego Gomez, Michael Bowling, Marlos C. MachadoICLR 2024 · 被引用 11 次
- Scalable Multi-agent Covering Option Discovery based on Kronecker GraphsJiayu Chen, Jingdi Chen, Tian Lan, Vaneet AggarwalNeurIPS 2022 · 被引用 16 次
- Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph DrawingKaixin Wang, Kuangqi Zhou, Qixin Zhang, Jie Shao 等ICML 2021 · 被引用 32 次
