Learning from Active Human Involvement through Proxy Value Propagation
Zhenghao Mark Peng, Wenjie Mo, Chenda Duan, Quanyi Li, Bolei Zhou
Abstract
Learning from active human involvement enables the human subject to actively intervene and demonstrate to the AI agent during training. The interaction and corrective feedback from human brings safety and AI alignment to the learning process. In this work, we propose a new reward-free active human involvement method called Proxy Value Propagation for policy optimization. Our key insight is that a proxy value function can be designed to express human intents, wherein state-action pairs in the human demonstration are labeled with high values, while those agents' actions that are intervened receive low values. Through the TD-learning framework, labeled values of demonstrated state-action pairs are further propagated to other unlabeled data generated from agents' exploration. The proxy value function thus induces a policy that faithfully emulates human behaviors. Human-in-the-loop experiments show the generality and efficiency of our method. With minimal modification to existing reinforcement learning algorithms, our method can learn to solve continuous and discrete control tasks with various human control devices, including the challenging task of driving in Grand Theft Auto V. Demo video and code are available at: https://metadriverse.github.io/pvp
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c3bc3df-8df8-485c-bad3-4f8099c136f6Cited by top-tier papers7
- Shared Autonomy with IDA: Interventional Diffusion AssistanceBrandon McMahan, Zhenghao Mark Peng, Bolei Zhou, Jonathan C. KaoNeurIPS 2024 · 12 citations
- Predictive Preference Learning from Human InterventionsHaoyuan Cai, Zhenghao Mark Peng, Bolei ZhouNeurIPS 2025 · 6 citations
- AURA: Multi-modal Shared Autonomy for Urban NavigationYukai Ma, Honglin He, Selina Song, Wayne Wu et al.CVPR 2026
- Policy Optimization under Imperfect Human Interactions with Agent-Gated Shared AutonomyZhenghai Xue, Bo An, Shuicheng YanICLR 2025
- Faithful Dynamic Imitation Learning from Human Intervention with Dynamic Regret MinimizationBo Ling, Zhengyu Gan, Wanyuan Wang, Guanyu Gao et al.NeurIPS 2025
Builds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- DeepTake: Prediction of Driver Takeover Behavior using Multimodal DataErfan Pakdamanian, Shili Sheng, Sonia Baee, Seongkook Heo et al.CHI 2021 · 86 citations
Related papers
- Efficient Learning of Safe Driving Policy via Human-AI Copilot OptimizationQuanyi Li, Zhenghao Peng, Bolei ZhouICLR 2022 · 80 citations
- Reinforcement Learning from Imperfect Corrective Actions and Proxy RewardsZhaohui Jiang, Xuening Feng, Paul Weng, Yifei Zhu et al.ICLR 2025
- Robot-Gated Interactive Imitation Learning with Adaptive Intervention MechanismHaoyuan Cai, Zhenghao Peng, Bolei ZhouICML 2025
- What about Inputting Policy in Value Function: Policy Representation and Policy-Extended Value Function ApproximatorHongyao Tang, Zhaopeng Meng, Jianye Hao, Chen Chen et al.AAAI 2022 · 20 citations
- Human-AI Shared Control via Policy DissectionQuanyi Li, Zhenghao Peng, Haibin Wu, Lan Feng et al.NeurIPS 2022 · 16 citations
