Uncertainty-Aware Policy Optimization: A Robust, Adaptive Trust Region Approach
James Queeney, Ioannis Ch. Paschalidis, Christos G. Cassandras
摘要
In order for reinforcement learning techniques to be useful in real-world decision making processes, they must be able to produce robust performance from limited data. Deep policy optimization methods have achieved impressive results on complex tasks, but their real-world adoption remains limited because they often require significant amounts of data to succeed. When combined with small sample sizes, these methods can result in unstable learning due to their reliance on high-dimensional sample-based estimates. In this work, we develop techniques to control the uncertainty introduced by these estimates. We leverage these techniques to propose a deep policy optimization approach designed to produce stable performance even when data is scarce. The resulting algorithm, Uncertainty-Aware Trust Region Policy Optimization, generates robust policy updates that adapt to the level of uncertainty present throughout the learning process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DFF: Decision-Focused Fine-Tuning for Smarter Predict-Then-Optimize with Limited DataJiaqi Yang, Enming Liang, Zicheng Su, Zhichao Zou 等AAAI 2025 · 被引用 6 次
- Predictive CVaR Q-learningJu-Hyun Kim, Seungki MinICLR 2026
相关 Paper
- Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous ControlNate Rahn, Pierluca D'Oro, Harley Wiltzer, Pierre-Luc Bacon 等NeurIPS 2023 · 被引用 11 次
- Monotonic Robust Policy Optimization with Model DiscrepancyYuankun Jiang, Chenglin Li, Wenrui Dai, Junni Zou 等ICML 2021 · 被引用 24 次
- Learning Robust Policy against Disturbance in Transition Dynamics via State-Conservative Policy OptimizationYufei Kuang, Miao Lu, Jie Wang, Qi Zhou 等AAAI 2022 · 被引用 29 次
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 305 次
- Deep Bayesian Quadrature Policy OptimizationRavi Tej Akella, Kamyar Azizzadenesheli, Mohammad Ghavamzadeh, Animashree Anandkumar 等AAAI 2021 · 被引用 5 次
