Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access
Daniel Ebi, Damien Ernst, Klemens Böhm, Gaspard Lambrechts
Abstract
Asymmetric reinforcement learning leverages privileged information available during training to improve learning under partial observability. Existing asymmetric actor-critic methods typically assume access to the full environment state to condition the critic during training, which is often unrealistic in practice. We introduce the informed asymmetric actor-critic framework that allows the critic to be conditioned on arbitrary state-dependent privileged signals, and show that any such signal yields unbiased policy gradient estimates. This substantially expands the set of admissible privileged information and raises the problem of selecting the most informative signals for learning. To this end, we propose two novel informativeness criteria: a dependence-based test that can be applied prior to training, and a test based on improvements in value prediction that can be applied post hoc. Experiments on partially observable benchmarks and synthetic environments demonstrate that carefully selected privileged signals can match or outperform full-state asymmetric baselines while relying on strictly less state information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 113841c4-221f-40ff-ab23-6b33fe6ecf36Builds on12
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 437 citations
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 270 citations
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement LearningZhaohan Daniel Guo, Bernardo Ávila Pires, Bilal Piot, Jean-Bastien Grill et al.ICML 2020 · 153 citations
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 75 citations
- Bridging State and History Representations: Understanding Self-Predictive RLTianwei Ni, Benjamin Eysenbach, Erfan Seyedsalehi, Michel Ma et al.ICLR 2024 · 50 citations
Related papers
- A Theoretical Justification for Asymmetric Actor-Critic AlgorithmsGaspard Lambrechts, Damien Ernst, Aditya MahajanICML 2025
- PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement LearningDongchi Huang, Jiaqi Wang, Yang Li, Chunhe Xia et al.ICML 2025
- Provable Partially Observable Reinforcement Learning with Privileged InformationYang Cai, Xiangyu Liu, Argyris Oikonomou, Kaiqing ZhangNeurIPS 2024 · 22 citations
- Robust Asymmetric Learning in POMDPsAndrew Warrington, Jonathan Wilder Lavington, Adam Scibior, Mark Schmidt et al.ICML 2021 · 37 citations
- Posterior Value Functions: Hindsight Baselines for Policy Gradient MethodsChris Nota, Philip S. Thomas, Bruno C. da SilvaICML 2021 · 6 citations
