Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
Hiroki Furuta, Tatsuya Matsushima, Tadashi Kozuno, Yutaka Matsuo, Sergey Levine, Ofir Nachum, Shixiang Shane Gu
Abstract
Progress in deep reinforcement learning (RL) research is largely enabled by benchmark task environments. However, analyzing the nature of those environments is often overlooked. In particular, we still do not have agreeable ways to measure the difficulty or solvability of a task, given that each has fundamentally different actions, observations, dynamics, rewards, and can be tackled with diverse RL algorithms. In this work, we propose policy information capacity (PIC) -- the mutual information between policy parameters and episodic return -- and policy-optimal information capacity (POIC) -- between policy parameters and episodic optimality -- as two environment-agnostic, algorithm-agnostic quantitative metrics for task difficulty. Evaluating our metrics across toy environments as well as continuous control benchmark tasks from OpenAI Gym and DeepMind Control Suite, we empirically demonstrate that these information-theoretic metrics have higher correlations with normalized task solvability scores than a variety of alternatives. Lastly, we show that these metrics can also be used for fast and compute-efficient optimizations of key design parameters such as reward shaping, policy architectures, and MDP properties for better solvability by RL algorithms without ever running full RL experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ae3d93d-af49-41b9-a5df-a34e272d64b7Cited by top-tier papers5
- Generalized Decision Transformer for Offline Hindsight Information MatchingHiroki Furuta, Yutaka Matsuo, Shixiang Shane GuICLR 2022 · 125 citations
- Explicable Reward Design for Reinforcement Learning AgentsRati Devidze, Goran Radanovic, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2021 · 60 citations
- Cliff Diving: Exploring Reward Surfaces in Reinforcement Learning EnvironmentsRyan Sullivan, J. K. Terry, Benjamin Black, John P. DickersonICML 2022 · 11 citations
- CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous QueriesNi Mu, Hao Hu, Xiao Hu, Yiqin Yang et al.ICML 2025
- A System for Morphology-Task Generalization via Unified Representation and Behavior DistillationHiroki Furuta, Yusuke Iwasawa, Yutaka Matsuo, Shixiang Shane GuICLR 2023
Builds on9
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering SkillsVictor Campos, Alexander Trott, Caiming Xiong, Richard Socher et al.ICML 2020 · 178 citations
Related papers
- Model-agnostic Measure of Generalization DifficultyAkhilan Boopathy, Kevin Liu, Jaedong Hwang, Shu Ge et al.ICML 2023 · 8 citations
- Policy-Independent Behavioral Metric-Based Representation for Deep Reinforcement LearningWeijian Liao, Zongzhang Zhang, Yang YuAAAI 2023 · 7 citations
- PMIC: Improving Multi-Agent Reinforcement Learning with Progressive Mutual Information CollaborationPengyi Li, Hongyao Tang, Tianpei Yang, Xiaotian Hao et al.ICML 2022 · 49 citations
- Predictive Information Accelerates Learning in RLKuang-Huei Lee, Ian Fischer, Anthony Z. Liu, Yijie Guo et al.NeurIPS 2020 · 82 citations
- APS: Active Pretraining with Successor FeaturesHao Liu, Pieter AbbeelICML 2021 · 147 citations
