Continual Learning for Instruction Following from Realtime Feedback
Alane Suhr, Yoav Artzi
Abstract
We propose and deploy an approach to continually train an instruction-following agent from feedback provided by users during collaborative interactions. During interaction, human users instruct an agent using natural language, and provide realtime binary feedback as they observe the agent following their instructions. We design a contextual bandit learning approach, converting user feedback to immediate reward. We evaluate through thousands of human-agent interactions, demonstrating 15.4% absolute improvement in instruction execution accuracy over time. We also show our approach is robust to several design variations, and that the feedback signal is roughly equivalent to the learning signal of supervised demonstration data. * Work done while at Cornell University. 2 We use the term continual learning to refer to a learning setting where an agent continually improves its task performance [14], in our case following instructions. The term is used at times to describe the domain adaptation challenge of continually learning to handle new tasks. Our work does not address this problem. 3 Throughout the paper, we use agent to refer to an automated instruction-following system. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35970d48-7a6b-4ba5-b4b9-3d950ff03f2bCited by top-tier papers6
- Bisecle: Binding and Separation in Continual Learning for Video Language UnderstandingYue Tan, Xiaoqian Hu, Hao Xue, Celso de Melo et al.NeurIPS 2025 · 14 citations
- φ-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal ModelsThanh-Dat Truong, Huu-Thien Tran, Jackson David Cothren, Bhiksha Raj et al.CVPR 2026 · 2 citations
- Global Reward to Local Rewards: Multimodal-Guided Decomposition for Improving Dialogue AgentsDong Won Lee, Hae Park, Yoon Kim, Cynthia Breazeal et al.EMNLP 2024 · 1 citation
- Grounding Language in Multi-Perspective Referential CommunicationZineng Tang, Lingjun Mao, Alane SuhrEMNLP 2024 · 1 citation
- Retrospective Learning from InteractionsZizhao Chen, Mustafa Omer Gul, Yiwei Chen, Gloria Geng et al.ACL 2025
Builds on5
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
- Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy OptimizationRajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel et al.ICLR 2023 · 54 citations
- An Imitation Game for Learning Semantic Parsers from User InteractionZiyu Yao, Yiqi Tang, Wen-tau Yih, Huan Sun et al.EMNLP 2020 · 18 citations
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk et al.CVPR 2020
- Simulating Bandit Learning from User Feedback for Extractive Question AnsweringGe Gao, Eunsol Choi, Yoav ArtziACL 2022
Related papers
- Continually Improving Extractive QA via Human FeedbackGe Gao, Hung-Ting Chen, Yoav Artzi, Eunsol ChoiEMNLP 2023 · 5 citations
- ConTinTin: Continual Learning from Task InstructionsWenpeng Yin, Jia Li, Caiming XiongACL 2022
- How to talk so AI will learn: Instructions, descriptions, and autonomyTheodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths et al.NeurIPS 2022 · 30 citations
- Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedbackJing Xu, Megan Ung, Mojtaba Komeili, Kushal Arora et al.ACL 2023 · 13 citations
- Just Ask: An Interactive Learning Framework for Vision and Language NavigationTa-Chung Chi, Minmin Shen, Mihail Eric, Seokhwan Kim et al.AAAI 2020 · 88 citations
