Continually Improving Extractive QA via Human Feedback
Ge Gao, Hung-Ting Chen, Yoav Artzi, Eunsol Choi
Abstract
We study continually improving an extractive question answering (QA) system via human user feedback. We design and deploy an iterative approach, where information-seeking users ask questions, receive model-predicted answers, and provide feedback. We conduct experiments involving thousands of user interactions under diverse setups to broaden the understanding of learning from feedback over time. Our experiments show effective improvement from user feedback of extractive QA models over time across different data regimes, including significant potential for domain adaptation. * Equal contribution. 1 The term continual learning is at times used to refer to a scenario where models adapt to new tasks over time. We study improving the model continually on its original task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c227e4e3-65ab-4e15-aefc-54a208683935Cited by top-tier papers4
- I Could've Asked That: Reformulating Unanswerable QuestionsWenting Zhao, Ge Gao, Claire Cardie, Alexander M. RushEMNLP 2024 · 2 citations
- CoGen: Learning from Feedback with Coupled Comprehension and GenerationMustafa Omer Gul, Yoav ArtziEMNLP 2024 · 1 citation
- Retrospective Learning from InteractionsZizhao Chen, Mustafa Omer Gul, Yiwei Chen, Gloria Geng et al.ACL 2025
- FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language ModelsJavier Carnerero-Cano, Massimiliano Pronesti, Radu Marinescu, Tigran T. Tchrakian et al.ACL 2026
Builds on5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedbackJing Xu, Megan Ung, Mojtaba Komeili, Kushal Arora et al.ACL 2023 · 13 citations
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson et al.EMNLP 2020 · 9 citations
- Simulating Bandit Learning from User Feedback for Extractive Question AnsweringGe Gao, Eunsol Choi, Yoav ArtziACL 2022
Related papers
- Continual Learning for Instruction Following from Realtime FeedbackAlane Suhr, Yoav ArtziNeurIPS 2023 · 27 citations
- Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System ImprovementBhavana Dalvi Mishra, Oyvind Tafjord, Peter ClarkEMNLP 2022 · 10 citations
- Multi-Source Test-Time Adaptation as Dueling Bandits for Extractive Question AnsweringHai Ye, Qizhe Xie, Hwee Tou NgACL 2023 · 2 citations
- Learning a Cost-Effective Annotation Policy for Question AnsweringBernhard Kratzwald, Stefan Feuerriegel, Huan SunEMNLP 2020 · 9 citations
- Continual Dialogue State Tracking via Example-Guided Question AnsweringHyundong Cho, Andrea Madotto, Zhaojiang Lin, Khyathi Raghavi Chandu et al.EMNLP 2023
