Reinforcement Learning Control of a Physical Robot Device for Assisted Human Walking without a Simulator
Junmin Zhong, Emiliano Quiñones Yumbla, Seyed Yousef Soltanian, Ruofan Wu, Wenlong Zhang, Jennie Si
摘要
This study presents an innovative reinforcement learning (RL) control approach to facilitate soft exosuit-assisted human walking. Our goal is to address the ongoing challenges in developing reliable RL-based methods for controlling physical devices. To overcome key obstacles-such as limited data, the absence of a simulator for human-robot interaction during walking, the need for low computational overhead in real-time deployment, and the demand for rapid adaptation to achieve personalized control while ensuring human safety-we propose an online Adaptation from an offline Imitating Expert Policy (AIP) approach. Our offline learning mimics human expert actions through real human walking demonstrations without robot assistance. The resulted policy is then used to initialize online actor-critic learning, the goal of which is to optimally personalize robot assistance. In addition to being fast and robust, our online RL method also posses important properties such as learning convergence, dynamic stability, and solution optimality. We have successfully demonstrated our simple and robust framework for safe robot control on all five tested human participants, without selectively presenting results. The qualitative performance guarantees provided by our online RL, along with the consistent experimental validation of AIP control, represent the first demonstration of online adaptation for softsuit control personalization and serve as important evidence for the use of online RL in controlling a physical device to solve a real-life problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Data Quality in Imitation LearningSuneel Belkhale, Yuchen Cui, Dorsa SadighNeurIPS 2023 · 被引用 135 次
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li 等NeurIPS 2022 · 被引用 81 次
- Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy OptimizationKun Lei, Zhengmao He, Chenhao Lu, Kaizhe Hu 等ICLR 2024 · 被引用 31 次
相关 Paper
- Exo-Plore: Exploring Exoskeleton Control Space through Human-aligned SimulationGeonho Leem, Jaedong Lee, Jehee Lee, Seungmoon Song 等ICLR 2026 · 被引用 11 次
- Human-Robotic Prosthesis as Collaborating Agents for Symmetrical WalkingRuofan Wu, Junmin Zhong, Brent Wallace, Xiang Gao 等NeurIPS 2022 · 被引用 16 次
- ReActor: Reinforcement Learning for Physics-Aware Motion RetargetingDavid Müller, Agon Serifi, Sammy Christen, Ruben Grandia 等SIGGRAPH 2026
- Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid ControlWeidong Huang, Zhehan Li, Hangxin Liu, Biao Hou 等ICLR 2026 · 被引用 4 次
- Fine-tuning Behavioral Cloning Policies with Preference‑Based Reinforcement LearningMaël Macuglia, Paul Friedrich, Giorgia RamponiICLR 2026 · 被引用 2 次
