: On-Device Real-Time Deep Reinforcement Learning for Autonomous Robotics
Zexin Li, Aritra Samanta, Yufei Li, Andrea Soltoggio, Hyoseung Kim, Cong Liu
摘要
Autonomous robotic systems, like autonomous vehicles and robotic search and rescue, require efficient on-device training for continuous adaptation of Deep Reinforcement Learning (DRL) models in dynamic environments. This research is fundamentally motivated by the need to understand and address the challenges of on-device real-time DRL, which involves balancing timing and algorithm performance under memory constraints, as exposed through our extensive empirical studies. This intricate balance requires co-optimizing two pivotal parameters of DRL training -batch size and replay buffer size. Configuring these parameters significantly affects timing and algorithm performance, while both (unfortunately) require substantial memory allocation to achieve near-optimal performance. This paper presents R 3 , a holistic solution for managing timing, memory, and algorithm performance in on-device real-time DRL training. R 3 employs (i) a deadline-driven feedback loop with dynamic batch sizing for optimizing timing, (ii) efficient memory management to reduce memory footprint and allow larger replay buffer sizes, and (iii) a runtime coordinator guided by heuristic analysis and a runtime profiler for dynamically adjusting memory resource reservations. These components collaboratively tackle the trade-offs in on-device DRL training, improving timing and algorithm performance while minimizing the risk of out-ofmemory (OOM) errors.
We implemented and evaluated R 3 extensively across various DRL frameworks and benchmarks on three hardware platforms commonly adopted by autonomous robotic systems. Additionally, we integrate R 3 with a popular realistic autonomous car simulator to demonstrate its real-world applicability. Evaluation results show that R 3 achieves efficacy across diverse platforms, ensuring consistent latency performance and timing predictability with minimal overhead. Moreover, R 3 showcases versatility by handling varied optimization goals and adapting to fluctuating systems scenarios.
1 "Algorithm performance" in DRL is akin to "accuracy" in DNNs. In DNNs, accuracy gauges the correct prediction rate. However, in DRL, performance is more multifaceted, primarily involving the agent's capability to optimize cumulative reward over time, balancing exploration and exploitation. While accuracy can be relevant in some tasks, DRL performance typically entails a broader, more complex set of considerations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio 等ICML 2020 · 被引用 303 次
- Response Time Analysis and Priority Assignment of Processing Chains on ROS2 ExecutorsYue Tang, Zhiwei Feng, Nan Guan, Xu Jiang 等RTSS 2020 · 被引用 81 次
- LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN TasksWoosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin 等RTSS 2021 · 被引用 68 次
- A ROS 2 Response-Time Analysis Exploiting Starvation Freedom and Execution-Time VarianceTobias Blaß, Daniel Casini, Sergey Bozhko, Björn B. BrandenburgRTSS 2021 · 被引用 64 次
- End-To-End Timing Analysis in ROS2Harun Teper, Mario Günzel, Niklas Ueter, Georg von der Brüggen 等RTSS 2022 · 被引用 56 次
相关 Paper
- DuoJoule: Accurate On-Device Deep Reinforcement Learning for Energy and TimelinessSoheil Shirvani, Aritra Samanta, Zexin Li, Cong LiuRTSS 2024 · 被引用 2 次
- A3C-S: Automated Agent Accelerator Co-Search towards Efficient Deep Reinforcement LearningYonggan Fu, Yongan Zhang, Chaojian Li, Zhongzhi Yu 等DAC 2021 · 被引用 4 次
- Adjustable Robust Reinforcement Learning for Online 3D Bin PackingYuxin Pan, Yize Chen, Fangzhen LinNeurIPS 2023 · 被引用 23 次
- Sample-Efficient Automated Deep Reinforcement LearningJörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank HutterICLR 2021 · 被引用 49 次
- Simple Augmentation Goes a Long Way: ADRL for DNN QuantizationLin Ning, Guoyang Chen, Weifeng Zhang, Xipeng ShenICLR 2021 · 被引用 1 次
