Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System Improvement
Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Clark
Abstract
Our goal is a teachable reasoning system for question-answering (QA), where a user can interact with faithful answer explanations, and correct its errors so that the system improves over time. Our approach is to augment a QA model with a dynamic memory of user feedback, containing user-supplied corrections to erroneous model beliefs that users identify during interaction. Retrievals from memory are used as additional context for QA, to help avoid previous mistakes in similar new situationsa novel application of memory-based continuous learning. With simulated feedback, we find that our system (called TeachMe 1 ) continually improves with time, and without model retraining, requiring feedback on only 25% of training examples to reach within 1% of the upper-bound (feedback on all examples). Similarly, in experiments with real users, we observe a similar trend, with performance improving by over 15% on a hidden test set after teaching. This suggests new opportunities for using frozen language models in an interactive setting where users can inspect, debug, and correct the model's beliefs, leading to improved system's performance over time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43d2b398-e534-4b2f-be66-d46a3f17ad04Cited by top-tier papers8
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
- Learning Deductive Reasoning from Synthetic Corpus based on Formal LogicTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaICML 2023 · 45 citations
- Entailer: Answering Questions with Faithful and Truthful Chains of ReasoningOyvind Tafjord, Bhavana Dalvi Mishra, Peter ClarkEMNLP 2022 · 28 citations
- How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent AdvancesZihan Zhang, Meng Fang, Ling Chen, Mohammad-Reza Namazi-Rad et al.EMNLP 2023 · 22 citations
- Faithful Question Answering with Monte-Carlo PlanningRuixin Hong, Hongming Zhang, Hong Zhao, Dong Yu et al.ACL 2023 · 7 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
- A Rationale-Centric Framework for Human-in-the-loop Machine LearningJinghui Lu, Linyi Yang, Brian MacNamee, Yue ZhangACL 2022 · 46 citations
Related papers
- Continually Improving Extractive QA via Human FeedbackGe Gao, Hung-Ting Chen, Yoav Artzi, Eunsol ChoiEMNLP 2023 · 5 citations
- Memory-assisted prompt editing to improve GPT-3 after deploymentAman Madaan, Niket Tandon, Peter Clark, Yiming YangEMNLP 2022 · 1 citation
- Remembering for the Right Reasons: Explanations Reduce Catastrophic ForgettingSayna Ebrahimi, Suzanne Petryk, Akash Gokul, William Gan et al.ICLR 2021 · 3 citations
- MAF: Multi-Aspect Feedback for Improving Reasoning in Large Language ModelsDeepak Nathani, David Wang, Liangming Pan, William Yang WangEMNLP 2023 · 6 citations
- BeliefBank: Adding Memory to a Pre-Trained Language Model for a Systematic Notion of BeliefNora Kassner, Oyvind Tafjord, Hinrich Schütze, Peter ClarkEMNLP 2021 · 2 citations
