Learning Interpretable Options by Identifying Reward Diffusion Bottlenecks in Reinforcement Learning
Yiming Fei, Ziming Wang, Rui Yan, Huajin Tang
Abstract
Bottleneck states, which connect distinct regions of the state space, provide a principled and interpretable basis for constructing temporal abstractions in Hierarchical Reinforcement Learning (HRL). However, existing bottleneck identification methods primarily rely on topological analysis of the state-transition graph, limiting their scalability to high-dimensional or continuous domains. To address this challenge, we introduce Value Power Strength (VPS), a value function-based metric inspired by the analogy between the Bellman equation and Kirchhoff’s current law, to quantify bottleneck property via the diffusion of reward in Markov Decision Processes (MDPs). VPS is estimated efficiently using value functions learned from random reward signals and captures reward diffusion bottlenecks in both discrete and continuous state spaces. Leveraging VPS, we design options that guide agents toward or away from bottleneck regions. Experimental results on classic tabular domains, continuous-control PointMaze, and Atari 2600 games demonstrate that the VPS-based framework discovers semantically meaningful subgoals and substantially improves exploration efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47c713e0-2cda-49a2-adb9-d81a9f004e63Builds on10
- Towards Better Laplacian Representation in Reinforcement Learning with Generalized Graph DrawingKaixin Wang, Kuangqi Zhou, Qixin Zhang, Jie Shao et al.ICML 2021 · 32 citations
- Deep Laplacian-based Options for Temporally-Extended ExplorationMartin Klissarov, Marlos C. MachadoICML 2023 · 31 citations
- Exploring through Random Curiosity with General Value FunctionsAditya A. Ramesh, Louis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberNeurIPS 2022 · 13 citations
- Proper Laplacian Representation LearningDiego Gomez, Michael Bowling, Marlos C. MachadoICLR 2024 · 11 citations
- Effectively Learning Initiation Sets in Hierarchical Reinforcement LearningAkhil Bagaria, Ben Abbatematteo, Omer Gottesman, Matt Corsaro et al.NeurIPS 2023 · 9 citations
Related papers
- Learning Subgoal Representations with Slow DynamicsSiyuan Li, Lulu Zheng, Jianhao Wang, Chongjie ZhangICLR 2021 · 48 citations
- Hierarchical Reinforcement Learning with Uncertainty-Guided Diffusional SubgoalsVivienne Huiling Wang, Tinghuai Wang, Joni PajarinenICML 2025
- Weakly-Supervised Reinforcement Learning for Controllable BehaviorLisa Lee, Ben Eysenbach, Ruslan Salakhutdinov, Shixiang Shane Gu et al.NeurIPS 2020 · 28 citations
- Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon ReasoningDhruv Shah, Peng Xu, Yao Lu, Ted Xiao et al.ICLR 2022 · 50 citations
- Graph-Theoretic Intrinsic Reward: Guiding RL with Effective ResistanceJatin Chauhan, Shivam Bhardwaj, Aditya Saibewar, Aditya Ramesh et al.ICLR 2026
