Learning from Mistakes via Cooperative Study Assistant for Large Language Models
Danqing Wang, Lei Li
Abstract
Large language models (LLMs) have demonstrated their potential to refine their generation based on their own feedback. However, the feedback from LLM itself is often inaccurate, thereby limiting its benefits. In this paper, we propose Study Assistant for Large LAnguage Model (SALAM), a novel framework with an auxiliary agent to assist the main LLM in learning from mistakes through interactive cooperation. In the gathering phase, the student assistant agent probes the main LLM, analyzes its errors, and collects the interaction in a mistake memory. During the examination phase, the study assistant provides guidelines by retrieving relevant cases to help the main LLM anticipate and avoid similar errors. We first investigate the effectiveness of a general study assistant and then customize it to provide LLMspecific guidance through imitation learning from successful guidance experiences. Our experiments on three LLMs using two challenging frameworks demonstrate that SALAM can significantly boost LLMs by an accuracy margin of up to 6.6 on BBH and 12.6 on BBQ 1 . 1 https://dqwang122.github.io/projects/SALAM . Jane thought today is 3/11/2002, but today is in fact Mar 12, which is 1 day later. What is the date a month ago? 02/11/2002 False Guideline: For dates in a problem, identify the correct date from which calculations should be made.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi et al.EMNLP 2024 · 119 citations
- Temporal Knowledge Question Answering via Abstract Reasoning InductionZiyang Chen, Dongfang Li, Xiang Zhao, Baotian Hu et al.ACL 2024 · 9 citations
- Mitigating Social Bias in Large Language Models: A Multi-Objective Approach Within a Multi-Agent FrameworkZhenjie Xu, Wenqing Chen, Yi Tang, Xuanying Li et al.AAAI 2025 · 5 citations
- QueryAgent: A Reliable and Efficient Reasoning Framework with Environmental Feedback based Self-CorrectionXiang Huang, Sitao Cheng, Shanshan Huang, Jiayu Shen et al.ACL 2024 · 3 citations
- Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale TuningSohan Patnaik, Milan Aggarwal, Sumit Bhatia, Balaji KrishnamurthyACL 2025 · 2 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
Related papers
- RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelYunda Tsai, Mingjie Liu, Haoxing RenDAC 2024 · 95 citations
- Generator-Assistant Stepwise Rollback Framework for Large Language Model AgentXingzuo Li, Kehai Chen, Yunfei Long, Xuefeng Bai et al.EMNLP 2025 · 1 citation
- Boosting the Potential of Large Language Models with an Intelligent Information AssistantYujia Zhou, Zheng Liu, Zhicheng DouNeurIPS 2024
- Experiential Co-Learning of Software-Developing AgentsChen Qian, Yufan Dang, Jiahao Li, Wei Liu et al.ACL 2024
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based AgentsTao Wu, Jingyuan Chen, Wang Lin, Mengze Li et al.ACL 2025 · 16 citations
