Real-DRL: Teach and Learn at Runtime
Yanbing Mao, Yihao Cai, Lui Sha
摘要
This paper introduces the Real-DRL framework for safety-critical autonomous systems, enabling runtime learning of a deep reinforcement learning (DRL) agent to develop safe and high-performance action policies in real plants (i.e., real physical systems to be controlled), while prioritizing safety! The Real-DRL consists of three interactive components: a DRL-Student, a PHY-Teacher, and a Trigger. The DRL-Student is a DRL agent that innovates in the dual self-learning and teaching-to-learn paradigm and the real-time safety-informed batch sampling. On the other hand, PHY-Teacher is a physics-model-based design of action policies that focuses solely on safety-critical functions. PHY-Teacher is novel in its real-time patch for two key missions: i) fostering the teaching-to-learn paradigm for DRL-Student and ii) backing up the safety of real plants. The Trigger manages the interaction between the DRL-Student and the PHY-Teacher. Powered by the three interactive components, the Real-DRL can effectively address safety challenges that arise from the unknown unknowns and the Sim2Real gap. Additionally, Real-DRL notably features i) assured safety, ii) automatic hierarchy learning (i.e., safety-first learning and then high-performance learning), and iii) safety-informed batch sampling to address the learning experience imbalance caused by corner cases. Experiments with a real quadruped robot, a quadruped robot in NVIDIA Isaac Gym, and a cart-pole system, along with comparisons and ablation studies, demonstrate the Real-DRL’s effectiveness and unique features.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Coupling Vision and Proprioception for Navigation of Legged RobotsZipeng Fu, Ashish Kumar, Ananye Agarwal, Haozhi Qi 等CVPR 2022 · 被引用 36 次
- Dynamic Model Predictive Shielding for Provably Safe Reinforcement LearningArko Banerjee, Kia Rahmani, Joydeep Biswas, Isil DilligNeurIPS 2024 · 被引用 24 次
- Posterior Sampling for Deep Reinforcement LearningRemo Sasso, Michelangelo Conserva, Paulo E. RauberICML 2023 · 被引用 14 次
- Physics-Regulated Deep Reinforcement Learning: Invariant EmbeddingsHongpeng Cao, Yanbing Mao, Lui Sha, Marco CaccamoICLR 2024 · 被引用 11 次
相关 Paper
- Catch Me If You Learn: Real-Time Attack Detection and Mitigation in Learning Enabled CPSIpsita Koley, Sunandan Adhikary, Soumyajit DeyRTSS 2021 · 被引用 8 次
- Conservative and Adaptive Penalty for Model-Based Safe Reinforcement LearningYecheng Jason Ma, Andrew Shen, Osbert Bastani, Dinesh JayaramanAAAI 2022 · 被引用 32 次
- Learning from the Test: Self-Referential Differential Testing for Deep RL AgentsJunda He, Jieke Shi, Zhou Yang, Mingfei Cheng 等ISSTA 2026
- Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real DeploymentGengyue Han, Yiheng FengICML 2026
- Addressing Action Oscillations through Learning Policy InertiaChen Chen, Hongyao Tang, Jianye Hao, Wulong Liu 等AAAI 2021 · 被引用 27 次
