Interaction-Grounded Learning
Tengyang Xie, John Langford, Paul Mineiro, Ida Momennejad
摘要
Consider a prosthetic arm, learning to adapt to its user's control signals. We propose Interaction-Grounded Learning for this novel setting, in which a learner's goal is to interact with the environment with no grounding or explicit reward to optimize its policies. Such a problem evades common RL solutions which require an explicit reward. The learning agent observes a multidimensional context vector, takes an action, and then observes a multidimensional feedback vector. This multidimensional feedback vector has no explicit reward information. In order to succeed, the algorithm must learn how to evaluate the feedback vector to discover a latent reward signal, with which it can ground its policies without supervision. We show that in an Interaction-Grounded Learning setting, with certain natural assumptions, a learner can discover the latent reward and ground its policy for successful interaction. We provide theoretical guarantees and a proof-of-concept empirical evaluation to demonstrate the effectiveness of our proposed approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Interaction-Grounded Learning with Action-Inclusive FeedbackTengyang Xie, Akanksha Saran, Dylan J. Foster, Lekan P. Molu 等NeurIPS 2022 · 被引用 12 次
- Formalizing Learning from Language Feedback with Provable GuaranteesWanqiao Xu, Allen Nie, Ruijie Zheng, Aditya Modi 等ICML 2026 · 被引用 8 次
- Learning to Guide and to be Guided in the Architect-Builder ProblemPaul Barde, Tristan Karch, Derek Nowrouzezahrai, Clément Moulin-Frier 等ICLR 2022 · 被引用 5 次
- Provably Efficient Interactive-Grounded Learning with Personalized RewardMengxiao Zhang, Yuheng Zhang, Haipeng Luo, Paul MineiroNeurIPS 2024 · 被引用 3 次
- An Information Theoretic Approach to Interaction-Grounded LearningXiaoyan Hu, Farzan Farnia, Ho-fung LeungICML 2024 · 被引用 3 次
它引用的顶会 Paper3
- Escaping the Gravitational Pull of SoftmaxJincheng Mei, Chenjun Xiao, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 56 次
- Latent Bandits RevisitedJoey Hong, Branislav Kveton, Manzil Zaheer, Yinlam Chow 等NeurIPS 2020 · 被引用 55 次
- Online Algorithm for Unsupervised Sequential Selection with Contextual InformationArun Verma, Manjesh Kumar Hanawal, Csaba Szepesvári, Venkatesh SaligramaNeurIPS 2020 · 被引用 6 次
相关 Paper
- Personalized Reward Learning with Interaction-Grounded Learning (IGL)Jessica Maghakian, Paul Mineiro, Kishan Panaganti, Mark Rucker 等ICLR 2023
- First Contact: Unsupervised Human-Machine Co-Adaptation via Mutual Information MaximizationSiddharth Reddy, Sergey Levine, Anca D. DraganNeurIPS 2022 · 被引用 18 次
- Learning Rewards From Linguistic FeedbackTheodore R. Sumers, Mark K. Ho, Robert X. D. Hawkins, Karthik Narasimhan 等AAAI 2021 · 被引用 67 次
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen 等ICML 2021 · 被引用 33 次
- Off-Policy Evaluation for Human FeedbackQitong Gao, Ge Gao, Juncheng Dong, Vahid Tarokh 等NeurIPS 2023 · 被引用 13 次
