Uncertainty-Sensitive Privileged Learning
Fan-Ming Luo, Lei Yuan, Yang Yu
摘要
Privileged learning efficiently tackles high-dimensional, partially observable decision-making problems by first training a privileged policy (PP) on lowdimensional privileged observations, and then deriving a deployment policy (DP) either by imitating the PP or coupling it with an observation encoder. However, since the DP relies on local and partial observations, a behavioral divergence (BD) often emerges between the DP and the PP, ultimately degrading deployment performance. A promising strategy is to train a PP to learn the optimal behaviors attainable under the DP's observation space by applying reward penalties in regions with large BD. However, producing these behaviors is challenging for the PP because they rely on the DP's information-gathering progress, which is invisible to the PP. In this paper, we quantify the DP's information-gathering progress by estimating the prediction uncertainty of privileged observations reconstructed from partial observations, and accordingly propose the framework of Uncertainty-Sensitive Privileged Learning (USPL). USPL feeds this uncertainty estimation to the PP and combines reward transformation with privileged-observation blurring, driving the PP to choose actions that actively reduce uncertainty and thus gather the necessary information. Experiments across nine tasks demonstrate that USPL significantly reduces the behavioral discrepancies, achieving superior deployment performance compared to baselines. Additional visualization results show that the DP accurately quantifies its uncertainty, and the PP effectively adapts to uncertainty variations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu 等ICML 2020 · 被引用 464 次
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 被引用 162 次
- When Do Transformers Shine in RL? Decoupling Memory from Credit AssignmentTianwei Ni, Michel Ma, Benjamin Eysenbach, Pierre-Luc BaconNeurIPS 2023 · 被引用 77 次
- Bridging the Imitation Gap by Adaptive InsubordinationLuca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador 等NeurIPS 2021 · 被引用 53 次
- Adapt to Environment Sudden Changes by Learning a Context Sensitive PolicyFan-Ming Luo, Shengyi Jiang, Yang Yu, Zongzhang Zhang 等AAAI 2022 · 被引用 40 次
相关 Paper
- Student-Informed Teacher TrainingNico Messikommer, Jiaxu Xing, Elie Aljalbout, Davide ScaramuzzaICLR 2025
- To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable RLYuda Song, Dhruv Rohatgi, Aarti Singh, J. Andrew BagnellNeurIPS 2025
- A Unified Principle of Pessimism for Offline Reinforcement Learning under Model MismatchYue Wang, Zhongchang Sun, Shaofeng ZouNeurIPS 2024 · 被引用 11 次
- Provable Partially Observable Reinforcement Learning with Privileged InformationYang Cai, Xiangyu Liu, Argyris Oikonomou, Kaiqing ZhangNeurIPS 2024 · 被引用 22 次
- PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement LearningDongchi Huang, Jiaqi Wang, Yang Li, Chunhe Xia 等ICML 2025
