First Contact: Unsupervised Human-Machine Co-Adaptation via Mutual Information Maximization
Siddharth Reddy, Sergey Levine, Anca D. Dragan
摘要
How can we train an assistive human-machine interface (e.g., an electromyographybased limb prosthesis) to translate a user's raw command signals into the actions of a robot or computer when there is no prior mapping, we cannot ask the user for supervision in the form of action labels or reward feedback, and we do not have prior knowledge of the tasks the user is trying to accomplish? The key idea in this paper is that, regardless of the task, when an interface is more intuitive, the user's commands are less noisy. We formalize this idea as a completely unsupervised objective for optimizing interfaces: the mutual information between the user's command signals and the induced state transitions in the environment. To evaluate whether this mutual information score can distinguish between effective and ineffective interfaces, we conduct a large-scale observational study on 540K examples of users operating various keyboard and eye gaze interfaces for typing, controlling simulated robots, and playing video games. The results show that our mutual information scores are predictive of the ground-truth task completion metrics in a variety of domains, with an average Spearman's rank correlation of ρ = 0.43. In addition to offline evaluation of existing interfaces, we use our unsupervised objective to learn an interface from scratch: we randomly initialize the interface, have the user attempt to perform their desired tasks using the interface, measure the mutual information score, and update the interface to maximize mutual information through reinforcement learning. We evaluate our method through a small-scale user study with 12 participants who perform a 2D cursor control task using a perturbed mouse, and an experiment with one expert user playing the Lunar Lander game using hand gestures captured by a webcam. The results show that we can learn an interface from scratch, without any user supervision or prior knowledge of tasks, with less than 30 minutes of human-in-the-loop training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Learning to Assist Humans without Inferring RewardsVivek Myers, Evan Ellis, Sergey Levine, Benjamin Eysenbach 等NeurIPS 2024 · 被引用 16 次
- Self-Calibrating BCIs: Ranking and Recovery of Mental Targets Without LabelsJonathan Grizou, Carlos de la Torre-Ortiz, Tuukka RuotsaloNeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper4
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley 等ICLR 2020 · 被引用 176 次
- Randomized Ensembled Double Q-Learning: Learning Fast Without a ModelXinyue Chen, Che Wang, Zijian Zhou, Keith W. RossICLR 2021 · 被引用 26 次
- X2T: Training an X-to-Text Typing Interface with Online Learning from User FeedbackJensen Gao, Siddharth Reddy, Glen Berseth, Nicholas Hardy 等ICLR 2021 · 被引用 10 次
相关 Paper
- Interaction-Grounded LearningTengyang Xie, John Langford, Paul Mineiro, Ida MomennejadICML 2021 · 被引用 3 次
- A data-driven approach for learning to control computersPeter Conway Humphreys, David Raposo, Tobias Pohlen, Gregory Thornton 等ICML 2022 · 被引用 124 次
- Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill DiscoveryKristian Hartikainen, Xinyang Geng, Tuomas Haarnoja, Sergey LevineICLR 2020 · 被引用 94 次
- Which Mutual-Information Representation Learning Objectives are Sufficient for Control?Kate Rakelly, Abhishek Gupta, Carlos Florensa, Sergey LevineNeurIPS 2021 · 被引用 44 次
- Learning to Draw Is Learning to See: Analyzing Eye Tracking Patterns for Assisted Observational DrawingFengqi Liu, Longji Huang, Zhengyu Huang, Zeyu WangSIGGRAPH 2025 · 被引用 1 次
