Effective Exploration Based on the Structural Information Principles
Xianghua Zeng, Hao Peng, Angsheng Li
Abstract
Traditional information theory provides a valuable foundation for Reinforcement Learning, particularly through representation learning and entropy maximization for agent exploration. However, existing methods primarily concentrate on modeling the uncertainty associated with RL's random variables, neglecting the inherent structure within the state and action spaces. In this paper, we propose a novel Structural Information principles-based Effective Exploration framework, namely SI2E. Structural mutual information between two variables is defined to address the single-variable limitation in structural information, and an innovative embedding principle is presented to capture dynamics-relevant state-action representations. The SI2E analyzes value differences in the agent's policy between state-action pairs and minimizes structural entropy to derive the hierarchical state-action structure, referred to as the encoding tree. Under this tree structure, value-conditional structural entropy is defined and maximized to design an intrinsic reward mechanism that avoids redundant transitions and promotes enhanced coverage in the state-action space. Theoretical connections are established between SI2E and classical information-theoretic methodologies, highlighting our framework's rationality and advantage. Comprehensive evaluations in the MiniGrid, MetaWorld, and DeepMind Control Suite benchmarks demonstrate that SI2E significantly outperforms state-of-the-art exploration baselines regarding final performance and sample efficiency, with maximum improvements of 37.63% and 60.25%, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14b02a19-e03a-4d7a-bb84-28dc0c50e69aCited by top-tier papers2
- SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language ModelsLinqin Wang, Yaping Liu, Zhengtao Yu, Shengxiang Gao et al.AAAI 2025 · 3 citations
- Robustness Evaluation of Graph-based News Detection Using Network Structural InformationXianghua Zeng, Hao Peng, Angsheng LiKDD 2025 · 1 citation
Builds on15
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann et al.ICML 2020 · 584 citations
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 258 citations
- State Entropy Maximization with Random Encoders for Efficient ExplorationYounggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee et al.ICML 2021 · 158 citations
- Structural Entropy Guided Graph Hierarchical PoolingJunran Wu, Xueyuan Chen, Ke Xu, Shangzhe LiICML 2022 · 113 citations
Related papers
- Novelty Search in Representational Space for Sample Efficient ExplorationRuo Yu Tao, Vincent François-Lavet, Joelle PineauNeurIPS 2020 · 53 citations
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen et al.ICML 2021 · 33 citations
- Accelerating Reinforcement Learning with Value-Conditional State Entropy ExplorationDongyoung Kim, Jinwoo Shin, Pieter Abbeel, Younggyo SeoNeurIPS 2023 · 34 citations
- Causality-Aware Efficient Exploration for Cooperative Multi-Agent Reinforcement LearningHongye Cao, Tianpei Yang, Fan Feng, Hammadi Rafik Ouariachi et al.AAAI 2026
- Information Prioritization through Empowerment in Visual Model-based RLHomanga Bharadhwaj, Mohammad Babaeizadeh, Dumitru Erhan, Sergey LevineICLR 2022 · 35 citations
