Locally Persistent Exploration in Continuous Control Tasks with Sparse Rewards
Susan Amin, Maziar Gomrokchi, Hossein Aboutalebi, Harsh Satija, Doina Precup
摘要
A major challenge in reinforcement learning is the design of exploration strategies, especially for environments with sparse reward structures and continuous state and action spaces. Intuitively, if the reinforcement signal is very scarce, the agent should rely on some form of short-term memory in order to cover its environment efficiently. We propose a new exploration method, based on two intuitions: (1) the choice of the next exploratory action should depend not only on the (Markovian) state of the environment, but also on the agent's trajectory so far, and (2) the agent should utilize a measure of spread in the state space to avoid getting stuck in a small region. Our method leverages concepts often used in statistical physics to provide explanations for the behavior of simplified (polymer) chains in order to generate persistent (locally self-avoiding) trajectories in state space. We discuss the theoretical properties of locally self-avoiding walks and their ability to provide a kind of short-term memory through a decaying temporal correlation within the trajectory. We provide empirical evaluations of our approach in a simulated 2D navigation task, as well as higher-dimensional MuJoCo continuous control locomotion tasks with sparse rewards.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- TAAC: Temporally Abstract Actor-Critic for Continuous ControlHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2021 · 被引用 30 次
- Generative Planning for Temporally Coordinated Exploration in Reinforcement LearningHaichao Zhang, Wei Xu, Haonan YuICLR 2022 · 被引用 12 次
- Subwords as Skills: Tokenization for Sparse-Reward Reinforcement LearningDavid Yunis, Justin Jung, Falcon Z. Dai, Matthew R. WalterNeurIPS 2024 · 被引用 5 次
- Simultaneously Updating All Persistence Values in Reinforcement LearningLuca Sabbioni, Luca Al Daire, Lorenzo Bisi, Alberto Maria Metelli 等AAAI 2023 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Memory Based Trajectory-conditioned Policies for Learning from Sparse RewardsYijie Guo, Jongwook Choi, Marcin Moczulski, Shengyu Feng 等NeurIPS 2020 · 被引用 36 次
- Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic RewardsFaisal Mohamed, Catherine Ji, Benjamin Eysenbach, Glen BersethICLR 2026 · 被引用 1 次
- Go Beyond Imagination: Maximizing Episodic Reachability with World ModelsYao Fu, Run Peng, Honglak LeeICML 2023 · 被引用 1 次
- Rank the Episodes: A Simple Approach for Exploration in Procedurally-Generated EnvironmentsDaochen Zha, Wenye Ma, Lei Yuan, Xia Hu 等ICLR 2021 · 被引用 47 次
- Generative Exploration and ExploitationJiechuan Jiang, Zongqing LuAAAI 2020 · 被引用 6 次
