Continual Learning for Instruction Following from Realtime Feedback
Alane Suhr, Yoav Artzi
摘要
We propose and deploy an approach to continually train an instruction-following agent from feedback provided by users during collaborative interactions. During interaction, human users instruct an agent using natural language, and provide realtime binary feedback as they observe the agent following their instructions. We design a contextual bandit learning approach, converting user feedback to immediate reward. We evaluate through thousands of human-agent interactions, demonstrating 15.4% absolute improvement in instruction execution accuracy over time. We also show our approach is robust to several design variations, and that the feedback signal is roughly equivalent to the learning signal of supervised demonstration data. * Work done while at Cornell University. 2 We use the term continual learning to refer to a learning setting where an agent continually improves its task performance [14], in our case following instructions. The term is used at times to describe the domain adaptation challenge of continually learning to handle new tasks. Our work does not address this problem. 3 Throughout the paper, we use agent to refer to an automated instruction-following system. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Bisecle: Binding and Separation in Continual Learning for Video Language UnderstandingYue Tan, Xiaoqian Hu, Hao Xue, Celso de Melo 等NeurIPS 2025 · 被引用 14 次
- φ-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal ModelsThanh-Dat Truong, Huu-Thien Tran, Jackson David Cothren, Bhiksha Raj 等CVPR 2026 · 被引用 2 次
- Global Reward to Local Rewards: Multimodal-Guided Decomposition for Improving Dialogue AgentsDong Won Lee, Hae Park, Yoon Kim, Cynthia Breazeal 等EMNLP 2024 · 被引用 1 次
- Grounding Language in Multi-Perspective Referential CommunicationZineng Tang, Lingjun Mao, Alane SuhrEMNLP 2024 · 被引用 1 次
- Retrospective Learning from InteractionsZizhao Chen, Mustafa Omer Gul, Yiwei Chen, Gloria Geng 等ACL 2025
它引用的顶会 Paper5
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
- Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy OptimizationRajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel 等ICLR 2023 · 被引用 54 次
- An Imitation Game for Learning Semantic Parsers from User InteractionZiyu Yao, Yiqi Tang, Wen-tau Yih, Huan Sun 等EMNLP 2020 · 被引用 18 次
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk 等CVPR 2020
- Simulating Bandit Learning from User Feedback for Extractive Question AnsweringGe Gao, Eunsol Choi, Yoav ArtziACL 2022
相关 Paper
- Continually Improving Extractive QA via Human FeedbackGe Gao, Hung-Ting Chen, Yoav Artzi, Eunsol ChoiEMNLP 2023 · 被引用 5 次
- ConTinTin: Continual Learning from Task InstructionsWenpeng Yin, Jia Li, Caiming XiongACL 2022
- How to talk so AI will learn: Instructions, descriptions, and autonomyTheodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths 等NeurIPS 2022 · 被引用 30 次
- Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedbackJing Xu, Megan Ung, Mojtaba Komeili, Kushal Arora 等ACL 2023 · 被引用 13 次
- Just Ask: An Interactive Learning Framework for Vision and Language NavigationTa-Chung Chi, Minmin Shen, Mihail Eric, Seokhwan Kim 等AAAI 2020 · 被引用 88 次
