Provable Rich Observation Reinforcement Learning with Combinatorial Latent States
Dipendra Misra, Qinghua Liu, Chi Jin, John Langford
Abstract
We propose a novel setting for reinforcement learning that combines two common real-world difficulties: presence of observations (such as camera images) and factored states (such as location of objects). In our setting, the agent receives observations generated stochastically from a "latent" factored state. These observations are "rich enough" to enable decoding of the latent state and remove partial observability concerns. Since the latent state is combinatorial, the size of state space is exponential in the number of latent factors. We create a learning algorithm FactoRL (Fact-o-Rel) for this setting, which uses noise-contrastive learning to identify latent structures in emission processes and discover a factorized state space. We derive polynomial sample complexity guarantees for FactoRL which polynomially depend upon the number factors, and very weakly depend on the size of the observation space. We also provide a guarantee of polynomial time complexity when given access to an efficient planning algorithm.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Provably Efficient Causal Model-Based Reinforcement Learning for Systematic GeneralizationMirco Mutti, Riccardo De Santi, Emanuele Rossi, Juan Felipe Calderón et al.AAAI 2023 · 17 citations
- Reinforcement Learning Under Latent Dynamics: Toward Statistical and Algorithmic ModularityPhilip Amortila, Dylan J. Foster, Nan Jiang, Akshay Krishnamurthy et al.NeurIPS 2024 · 6 citations
- Rich-Observation Reinforcement Learning with Continuous Latent DynamicsYuda Song, Lili Wu, Dylan J. Foster, Akshay KrishnamurthyICML 2024 · 2 citations
Related papers
- Provably Filtering Exogenous Distractors using Multistep Inverse DynamicsYonathan Efroni, Dipendra Misra, Akshay Krishnamurthy, Alekh Agarwal et al.ICLR 2022 · 38 citations
- Learning the Linear Quadratic Regulator from Nonlinear ObservationsZakaria Mhammedi, Dylan J. Foster, Max Simchowitz, Dipendra Misra et al.NeurIPS 2020 · 33 citations
- Towards Principled Representation Learning from Videos for Reinforcement LearningDipendra Misra, Akanksha Saran, Tengyang Xie, Alex Lamb et al.ICLR 2024 · 9 citations
- Learning World Models with Identifiable FactorizationYuren Liu, Biwei Huang, Zhengmao Zhu, Hong-Long Tian et al.NeurIPS 2023 · 32 citations
- Provably sample-efficient RL with side information about latent dynamicsYao Liu, Dipendra Misra, Miro Dudík, Robert E. SchapireNeurIPS 2022 · 2 citations
