Novel Exploration via Orthogonality
Andreas Theophilou, Özgür Simsek
摘要
Efficient exploration remains one of the most important open problems in reinforcement learning. Discovering novel states or transitions requires policies that efficiently direct the agent away from the regions of the state space that are already well explored. We introduce Novel Exploration via Orthogonality (NEO) , an approach that automatically uncovers not only which regions of the environment are novel but also how to reach them by leveraging Laplacian representations. NEO uses the eigenvectors of a modified graph Laplacian to induce gradient flows from states that are frequently visited (less novel) to states that are seldom visited (more novel). We show that NEO’s modified Laplacian yields eigenvectors whose extreme values align with the most novel regions of the state space. We provide bounds for the eigenvalues of the modified Laplacian; and we show that the smoothest eigenvectors with real eigenvalues below certain thresholds provide guaranteed gradients to novel states for both undirected and directed graphs. In an empirical evaluation in online, incremental settings, NEO outperformed related state-of-the-art approaches, including eigen-options and cover options, in a large collection of undirected and directed environments with varying connectivity structures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- NovelD: A Simple yet Effective Exploration CriterionTianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu 等NeurIPS 2021 · 被引用 106 次
- Exploration and Anti-Exploration with Distributional Random Network DistillationKai Yang, Jian Tao, Jiafei Lyu, Xiu LiICML 2024 · 被引用 37 次
- Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph DrawingKaixin Wang, Kuangqi Zhou, Qixin Zhang, Jie Shao 等ICML 2021 · 被引用 32 次
- Deep Laplacian-based Options for Temporally-Extended ExplorationMartin Klissarov, Marlos C. MachadoICML 2023 · 被引用 31 次
- Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement LearningSam Lobel, Akhil Bagaria, George KonidarisICML 2023 · 被引用 29 次
相关 Paper
- Exploration in Reinforcement Learning with Deep Covering OptionsYuu Jinnai, Jee Won Park, Marlos C. Machado, George Dimitri KonidarisICLR 2020 · 被引用 64 次
- Proper Laplacian Representation LearningDiego Gomez, Michael Bowling, Marlos C. MachadoICLR 2024 · 被引用 11 次
- Online Laplacian-Based Representation Learning in Reinforcement LearningMaheed H. Ahmed, Jayanth Bhargav, Mahsa GhasemiICML 2025
- Option Discovery in the Absence of Rewards with Manifold AnalysisAmitay Bar, Ronen Talmon, Ron MeirICML 2020 · 被引用 6 次
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
