Provable Partially Observable Reinforcement Learning with Privileged Information
Yang Cai, Xiangyu Liu, Argyris Oikonomou, Kaiqing Zhang
摘要
Partial observability of the underlying states generally presents significant challenges for reinforcement learning (RL). In practice, certain privileged information, e.g., the access to states from simulators, has been exploited in training and has achieved prominent empirical successes. To better understand the benefits of privileged information, we revisit and examine several simple and practically used paradigms in this setting. Specifically, we first formalize the empirical paradigm of expert distillation (also known as teacher-student learning), demonstrating its pitfall in finding near-optimal policies. We then identify a condition of the partially observable environment, the deterministic filter condition, under which expert distillation achieves sample and computational complexities that are both polynomial. Furthermore, we investigate another useful empirical paradigm of asymmetric actor-critic, and focus on the more challenging setting of observable partially observable Markov decision processes. We develop a belief-weighted asymmetric actor-critic algorithm with polynomial sample and quasi-polynomial computational complexities, in which one key component is a new provable oracle for learning belief states that preserve filter stability under a misspecified model, which may be of independent interest. Finally, we also investigate the provable efficiency of partially observable multi-agent RL (MARL) with privileged information. We develop algorithms featuring centralized-training-with-decentralized-execution, a popular framework in empirical MARL, with polynomial sample and (quasi-)polynomial computational complexities in both paradigms above. Compared with a few recent related theoretical studies, our focus is on understanding practically inspired algorithmic paradigms, without computationally intractable oracles.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Guided Policy Optimization under Partial ObservabilityYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 被引用 4 次
- Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State AccessDaniel Ebi, Damien Ernst, Klemens Böhm, Gaspard LambrechtsICML 2026 · 被引用 1 次
- Scalable Solutions to Zero-Sum Partially Observable Stochastic Games Through Belief Aggregation with Approximation GuaranteesKim Hammar, Tansu AlpcanAAAI 2026 · 被引用 1 次
- PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement LearningDongchi Huang, Jiaqi Wang, Yang Li, Chunhe Xia 等ICML 2025
- A Theoretical Justification for Asymmetric Actor-Critic AlgorithmsGaspard Lambrechts, Damien Ernst, Aditya MahajanICML 2025
它引用的顶会 Paper30
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 被引用 304 次
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 被引用 226 次
- Kinematic State Abstraction and Provably Efficient Rich-Observation Reinforcement LearningDipendra Misra, Mikael Henaff, Akshay Krishnamurthy, John LangfordICML 2020 · 被引用 158 次
- A Sharp Analysis of Model-based Reinforcement Learning with Self-PlayQinghua Liu, Tiancheng Yu, Yu Bai, Chi JinICML 2021 · 被引用 137 次
- Optimistic Policy Optimization with Bandit FeedbackLior Shani, Yonathan Efroni, Aviv Rosenberg, Shie MannorICML 2020 · 被引用 100 次
相关 Paper
- To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable RLYuda Song, Dhruv Rohatgi, Aarti Singh, J. Andrew BagnellNeurIPS 2025
- Partially Observable Multi-agent RL with (Quasi-)Efficiency: The Blessing of Information SharingXiangyu Liu, Kaiqing ZhangICML 2023
- Robust Asymmetric Learning in POMDPsAndrew Warrington, Jonathan Wilder Lavington, Adam Scibior, Mark Schmidt 等ICML 2021 · 被引用 37 次
- Privileged Information Distillation for Language ModelsEmiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, Alexandre Lacoste 等ICML 2026 · 被引用 61 次
- Breaking the Computational Barrier: Provably Efficient Actor–Critic for Low-Rank MDPsRuiquan Huang, Donghao Li, Yingbin LIANG, Jing YangICML 2026
