Simulating Bandit Learning from User Feedback for Extractive Question Answering
Ge Gao, Eunsol Choi, Yoav Artzi
Abstract
We study learning from user feedback for extractive question answering by simulating feedback using supervised data. We cast the problem as contextual bandit learning, and analyze the characteristics of several learning scenarios with focus on reducing data annotation. We show that systems initially trained on few examples can dramatically improve given feedback from users on model-predicted answers, and that one can use existing datasets to deploy systems in new domains without any annotation effort, but instead improving the system on-the-fly via user feedback. 040 simple binary feedback, and creates a contextual 041 bandit learning scenario (Auer et al., 2002; Lang-042 ford and Zhang, 2007). Figure 1 illustrates this 043 learning signal and its potential. 044 We simulate user feedback using several widely 045 used QA datasets, and use it as a bandit signal for 046 learning. We study the empirical characteristics 047 of the learning process, including its performance, 048 sensitivity to initial system performance, and trade-049 offs between online and offline learning. We also 050 simulate zero-annotation domain adaptation, where 051 we deploy a QA system trained from supervised 052 data in one domain and adapt it solely from user 053 feedback in a new domain. 104 c i , . . . , c j where i, j ∈ [1, n] and i ≤ j in the 105 context c as an answer. When relevant, we denote 106 π θ as a QA model parameterized by θ. 107 We formalize learning as a contextual bandit 108 process: at each time step t, the model is given 109 a question-context pair (q (t) , c(t) ), predicts an an-110 swer span ŷ, and receives a reward r (t) ∈ IR. 111 The learner's goal is to maximize the total reward 112 T t=1 r (t) . This formulation reflects a setup where, 113 given a question-context pair, the QA system inter-114 acts with users, who validate the model-predicted 115 answer in context, and provide feedback which is 116 mapped to a numerical reward. 117 Learning Algorithm We learn using policy gra-118 dient. Our learner is similar to REINFORCE (Sut-119 ton and Barto, 1998; Williams, 2004), but we use 120 arg max to predict answers instead of Monte Carlo 121 sampling from the model's output distribution. 3 122 We study online and offline learning, also re-123 ferred to as on-and off-policy. In online learning 124 (Algorithm 1), the model identity is maintained be-125 tween prediction and update; the parameter values 126 that are updated are the same that were used to gen-127 erate the output receiving reward. This entails that 128 a reward is only used once, to update the model 129 after observing it. In offline learning (Algorithm 2), 130 this relation between update and prediction does 131 not hold. The learner observes reward, often across 132 many examples, and may use it to update the model 133 many times, even after the parameters drifted arbi-134 trarily far from these that generated the prediction. 135 In practice, we observe reward for the entire length 136 of the simulation (T steps) and then update for 137 E epochs. The reward is re-weighted to provide 138 an unbiased estimation using inverse propensity 139 score (IPS; Horvitz and Thompson, 1952). We clip 140 the debiasing coefficient to avoid amplifying exam-141 ples with large coefficients (line 10, Algorithm 2). 142 In general, offline learning is easier to implement 143 because updating the model is not integrated with 144 its deployment. Offline learning also uses a train-145 ing loop that is similar to optimization practices in 146 supervised learning. This allows to iterate over the 147 data multiple times, albeit with the same feedback
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86495fd8-e79b-4f15-8fec-3d74ecf9c3ffCited by top-tier papers6
- Continual Learning for Instruction Following from Realtime FeedbackAlane Suhr, Yoav ArtziNeurIPS 2023 · 27 citations
- RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model OutputsAfra Feyza Akyürek, Ekin Akyürek, Ashwin Kalyan, Peter Clark et al.ACL 2023 · 24 citations
- Tapered Off-Policy REINFORCE - Stable and efficient reinforcement learning for large language modelsNicolas Le Roux, Marc G. Bellemare, Jonathan Lebensold, Arnaud Bergeron et al.NeurIPS 2025 · 5 citations
- Continually Improving Extractive QA via Human FeedbackGe Gao, Hung-Ting Chen, Yoav Artzi, Eunsol ChoiEMNLP 2023 · 5 citations
- When is Tree Search Useful for LLM Planning? It Depends on the DiscriminatorZiru Chen, Michael White, Raymond J. Mooney, Ali Payani et al.ACL 2024 · 3 citations
Builds on5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage RetrievalDevang Kulshreshtha, Robert Belfer, Iulian Vlad Serban, Siva ReddyEMNLP 2021 · 11 citations
- Learning a Cost-Effective Annotation Policy for Question AnsweringBernhard Kratzwald, Stefan Feuerriegel, Huan SunEMNLP 2020 · 9 citations
- Few-Shot Question Answering by Pretraining Span SelectionOri Ram, Yuval Kirstain, Jonathan Berant, Amir Globerson et al.ACL 2021
- Online Learning Meets Machine Translation Evaluation: Finding the Best Systems with the Least Human EffortVânia Mendonça, Ricardo Rei, Luísa Coheur, Alberto Sardinha et al.ACL 2021
Related papers
- Towards Domain Adaptive Neural Contextual BanditsZiyan Wang, Xiaoming Huo, Hao WangICLR 2025
- Multi-Source Test-Time Adaptation as Dueling Bandits for Extractive Question AnsweringHai Ye, Qizhe Xie, Hwee Tou NgACL 2023 · 2 citations
- Learning to Answer from Correct DemonstrationsNirmit Joshi, Gene Li, Siddharth Bhandari, Shiva Prasad Kasiviswanathan et al.ICLR 2026 · 6 citations
- Bandit Learning with Predicted Context: Regret Analysis and Selective Context QueryJianyi Yang, Shaolei RenINFOCOM 2021 · 8 citations
- Off-policy Bandits with Deficient SupportNoveen Sachdeva, Yi Su, Thorsten JoachimsKDD 2020 · 22 citations
