f-Policy Gradients: A General Framework for Goal-Conditioned RL using f-Divergences
Siddhant Agarwal, Ishan Durugkar, Peter Stone, Amy Zhang
摘要
Goal-Conditioned Reinforcement Learning (RL) problems often have access to sparse rewards where the agent receives a reward signal only when it has achieved the goal, making policy optimization a difficult problem. Several works augment this sparse reward with a learned dense reward function, but this can lead to sub-optimal policies if the reward is misaligned. Moreover, recent works have demonstrated that effective shaping rewards for a particular problem can depend on the underlying learning algorithm. This paper introduces a novel way to encourage exploration called -Policy Gradients, or -PG. -PG minimizes the f-divergence between the agent's state visitation distribution and the goal, which we show can lead to an optimal policy. We derive gradients for various f-divergences to optimize this objective. Our learning paradigm provides dense learning signals for exploration in sparse reward settings. We further introduce an entropy-regularized policy optimization objective, that we call -MaxEnt RL (or -MaxEnt RL) as a special case of our objective. We show that several metric-based shaping rewards like L2 can be used with -MaxEnt RL, providing a common ground to study such metric-based shaping rewards with efficient exploration. We find that -PG has better performance compared to standard policy gradient methods on a challenging gridworld as well as the Point Maze and FetchReach environments. More information on our website https://agarwalsiddhant10.github.io/projects/fpg.html.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Score Models for Offline Goal-Conditioned Reinforcement LearningHarshit Sikchi, Rohan Chitnis, Ahmed Touati, Alborz Geramifard 等ICLR 2024 · 被引用 16 次
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 被引用 14 次
- State Entropy Regularization for Robust Reinforcement LearningYonatan Ashlag, Uri Koren, Mirco Mutti, Esther Derman 等NeurIPS 2025 · 被引用 9 次
- RLZero: Direct Policy Inference from Language Without In-Domain SupervisionHarshit Sikchi, Siddhant Agarwal, Pranaya Jajoo, Samyak Parajuli 等NeurIPS 2025 · 被引用 8 次
- IOSTOM: Offline Imitation Learning from Observations via State Transition Occupancy MatchingQuang Anh Pham, Janaka Chathuranga Brahmanage, Tien Mai, Akshat KumarNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho 等NeurIPS 2021 · 被引用 107 次
- The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsSerena Booth, W. Bradley Knox, Julie Shah, Scott Niekum 等AAAI 2023 · 被引用 103 次
- C-Learning: Learning to Achieve Goals via Recursive ClassificationBenjamin Eysenbach, Ruslan Salakhutdinov, Sergey LevineICLR 2021 · 被引用 96 次
- Replacing Rewards with Examples: Example-Based Policy Search via Recursive ClassificationBen Eysenbach, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2021 · 被引用 56 次
相关 Paper
- Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled RegularizationSafwan Labbi, Daniil Tiapkin, Paul Mangold, Eric MoulinesICLR 2026 · 被引用 2 次
- Reward Shaping for Reinforcement Learning with An Assistant Reward AgentHaozhe Ma, Kuankuan Sima, Thanh Vinh Vo, Di Fu 等ICML 2024 · 被引用 34 次
- Goal-Conditioned Q-learning as Knowledge DistillationAlexander Levine, Soheil FeiziAAAI 2023 · 被引用 4 次
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair 等ICML 2020 · 被引用 303 次
- Planning Goals for ExplorationEdward S. Hu, Richard Chang, Oleh Rybkin, Dinesh JayaramanICLR 2023 · 被引用 152 次
