Continually Improving Extractive QA via Human Feedback
Ge Gao, Hung-Ting Chen, Yoav Artzi, Eunsol Choi
摘要
We study continually improving an extractive question answering (QA) system via human user feedback. We design and deploy an iterative approach, where information-seeking users ask questions, receive model-predicted answers, and provide feedback. We conduct experiments involving thousands of user interactions under diverse setups to broaden the understanding of learning from feedback over time. Our experiments show effective improvement from user feedback of extractive QA models over time across different data regimes, including significant potential for domain adaptation. * Equal contribution. 1 The term continual learning is at times used to refer to a scenario where models adapt to new tasks over time. We study improving the model continually on its original task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- I Could've Asked That: Reformulating Unanswerable QuestionsWenting Zhao, Ge Gao, Claire Cardie, Alexander M. RushEMNLP 2024 · 被引用 2 次
- CoGen: Learning from Feedback with Coupled Comprehension and GenerationMustafa Omer Gul, Yoav ArtziEMNLP 2024 · 被引用 1 次
- Retrospective Learning from InteractionsZizhao Chen, Mustafa Omer Gul, Yiwei Chen, Gloria Geng 等ACL 2025
- FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language ModelsJavier Carnerero-Cano, Massimiliano Pronesti, Radu Marinescu, Tigran T. Tchrakian 等ACL 2026
它引用的顶会 Paper5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedbackJing Xu, Megan Ung, Mojtaba Komeili, Kushal Arora 等ACL 2023 · 被引用 13 次
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson 等EMNLP 2020 · 被引用 9 次
- Simulating Bandit Learning from User Feedback for Extractive Question AnsweringGe Gao, Eunsol Choi, Yoav ArtziACL 2022
相关 Paper
- Continual Learning for Instruction Following from Realtime FeedbackAlane Suhr, Yoav ArtziNeurIPS 2023 · 被引用 27 次
- Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System ImprovementBhavana Dalvi Mishra, Oyvind Tafjord, Peter ClarkEMNLP 2022 · 被引用 10 次
- Multi-Source Test-Time Adaptation as Dueling Bandits for Extractive Question AnsweringHai Ye, Qizhe Xie, Hwee Tou NgACL 2023 · 被引用 2 次
- Learning a Cost-Effective Annotation Policy for Question AnsweringBernhard Kratzwald, Stefan Feuerriegel, Huan SunEMNLP 2020 · 被引用 9 次
- Continual Dialogue State Tracking via Example-Guided Question AnsweringHyundong Cho, Andrea Madotto, Zhaojiang Lin, Khyathi Raghavi Chandu 等EMNLP 2023
