Human-AI Shared Control via Policy Dissection
Quanyi Li, Zhenghao Peng, Haibin Wu, Lan Feng, Bolei Zhou
Abstract
Human-AI shared control allows human to interact and collaborate with autonomous agents to accomplish control tasks in complex environments. Previous Reinforcement Learning (RL) methods attempted goal-conditioned designs to achieve human-controllable policies at the cost of redesigning the reward function and training paradigm. Inspired by the neuroscience approach to investigate the motor cortex in primates, we develop a simple yet effective frequency-based approach called Policy Dissection to align the intermediate representation of the learned neural controller with the kinematic attributes of the agent behavior. Without modifying the neural controller or retraining the model, the proposed approach can convert a given RL-trained policy into a human-controllable policy. We evaluate the proposed approach on many RL tasks such as autonomous driving and locomotion. The experiments show that human-AI shared control system achieved by Policy Dissection in driving task can substantially improve the performance and safety in unseen traffic scenes. With human in the inference loop, the locomotion robots also exhibit versatile controllable motion skills even though they are only trained to move forward. Our results suggest the promising direction of implementing human-AI shared autonomy through interpreting the learned representation of the autonomous agents. Code and demo videos are available at https://metadriverse.github.io/policydissect . 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba23bf16-70b4-447b-a3ca-99b17b4a8c0cCited by top-tier papers2
- ACE: Off-Policy Actor-Critic with Causality-Aware Entropy RegularizationTianying Ji, Yongyuan Liang, Yan Zeng, Yu Luo et al.ICML 2024 · 20 citations
- Policy Optimization under Imperfect Human Interactions with Agent-Gated Shared AutonomyZhenghai Xue, Bo An, Shuicheng YanICLR 2025
Builds on9
- Benchmarking Deep Learning Interpretability in Time Series PredictionsAya Abdelsalam Ismail, Mohamed K. Gunady, Héctor Corrada Bravo, Soheil FeiziNeurIPS 2020 · 249 citations
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal TransformersRuihan Yang, Minghao Zhang, Nicklas Hansen, Huazhe Xu et al.ICLR 2022 · 146 citations
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement LearningAkanksha Atrey, Kaleigh Clary, David D. JensenICLR 2020 · 108 citations
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature AttributionNikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha et al.ICLR 2020 · 99 citations
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg et al.ICML 2020 · 81 citations
Related papers
- Efficient Learning of Safe Driving Policy via Human-AI Copilot OptimizationQuanyi Li, Zhenghao Peng, Bolei ZhouICLR 2022 · 80 citations
- Assistive Teaching of Motor Control Tasks to HumansMegha Srivastava, Erdem Biyik, Suvir Mirchandani, Noah D. Goodman et al.NeurIPS 2022 · 12 citations
- CrossLoco: Human Motion Driven Control of Legged Robots via Guided Unsupervised Reinforcement LearningTianyu Li, Hyunyoung Jung, Matthew C. Gombolay, Yong Kwon Cho et al.ICLR 2024 · 14 citations
- Learning from Active Human Involvement through Proxy Value PropagationZhenghao Mark Peng, Wenjie Mo, Chenda Duan, Quanyi Li et al.NeurIPS 2023 · 30 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
