Can Language Models Teach? Teacher Explanations Improve Student Performance via Personalization
Swarnadeep Saha, Peter Hase, Mohit Bansal
Abstract
A hallmark property of explainable AI models is the ability to teach other agents, communicating knowledge of how to perform a task. While Large Language Models (LLMs) perform complex reasoning by generating explanations for their predictions, it is unclear whether they also make good teachers for weaker agents. To address this, we consider a student-teacher framework between two LLM agents and study if, when, and how the teacher should intervene with natural language explanations to improve the student's performance. Since communication is expensive, we define a budget such that the teacher only communicates explanations for a fraction of the data, after which the student should perform well on its own. We decompose the teaching problem along four axes: (1) if teacher's test time intervention improve student predictions, (2) when it is worth explaining a data point, (3) how the teacher should personalize explanations to better teach the student, and (4) if teacher explanations also improve student performance on future unexplained data. We first show that teacher LLMs can indeed intervene on student reasoning to improve their performance. Next, inspired by the Theory of Mind abilities of effective teachers, we propose building two few-shot mental models of the student. The first model defines an Intervention Function that simulates the utility of an intervention, allowing the teacher to intervene when this utility is the highest and improving student performance at lower budgets. The second model enables the teacher to personalize explanations for a particular student and outperform unpersonalized teachers. We also demonstrate that in multi-turn interactions, teacher explanations generalize and learning from explained data improves student performance on future unexplained data. Finally, we also verify that misaligned teachers can lower student performance to random chance by intentionally misleading them. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language ModelsGuangyan Li, Yongqiang Tang, Wensheng ZhangICML 2024 · 11 citations
- Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information FlowJiaqi Bai, Hongcheng Guo, Zhongyuan Peng, Jian Yang et al.AAAI 2025 · 7 citations
- Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale TuningSohan Patnaik, Milan Aggarwal, Sumit Bhatia, Balaji KrishnamurthyACL 2025 · 2 citations
- Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model TutorsNico Daheim, Jakub Macina, Manu Kapur, Iryna Gurevych et al.EMNLP 2024 · 1 citation
- SlimLLM: Accurate Structured Pruning for Large Language ModelsJialong Guo, Xinghao Chen, Yehui Tang, Yunhe WangICML 2025
Builds on11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceGagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok et al.CHI 2021 · 713 citations
- Specializing Smaller Language Models towards Multi-Step ReasoningYao Fu, Hao Peng, Litu Ou, Ashish Sabharwal et al.ICML 2023 · 347 citations
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 216 citations
Related papers
- Democratizing Reasoning Ability: Tailored Learning from Large Language ModelZhaoyang Wang, Shaohan Huang, Yuxuan Liu, Jiahai Wang et al.EMNLP 2023 · 8 citations
- ThinkTuning: Instilling Cognitive Reflections without DistillationAswin RRV, Jacob Dineen, Divij Handa, Md Nayem Uddin et al.EMNLP 2025 · 9 citations
- Explainable Active Learning (XAL): Toward AI Explanations as Interfaces for Machine TeachersBhavya Ghai, Q. Vera Liao, Yunfeng Zhang, Rachel K. E. Bellamy et al.CSCW 2020 · 107 citations
- Probing to Refine: Reinforcement Distillation of LLM Reasoners via Explanatory InversionZhen Tan, Chengshuai Zhao, Song Wang, Jundong Li et al.ICLR 2026
- Teach2Eval: An Interaction-Driven LLMs Evaluation Method via Teaching EffectivenessYuhang Zhou, Xutian Chen, Yixin Cao, Yuchen Ni et al.ICLR 2026
