Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
Nico Daheim, Jakub Macina, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan
Abstract
Large language models (LLMs) present an opportunity to scale high-quality personalized education to all. A promising approach towards this means is to build dialog tutoring models that scaffold students' problem-solving. However, even though existing LLMs perform well in solving reasoning questions, they struggle to precisely detect student's errors and tailor their feedback to these errors. Inspired by realworld teaching practice where teachers identify student errors and customize their response based on them, we focus on verifying student solutions and show how grounding to such verification improves the overall quality of tutor response generation. We collect a dataset of 1K stepwise math reasoning chains with the first error step annotated by teachers. We show empirically that finding the mistake in a student solution is challenging for current models. We propose and evaluate several verifiers for detecting these errors. Using both automatic and human evaluation we show that the student solution verifiers steer the generation model towards highly targeted responses to student errors which are more often correct with less hallucinations compared to existing baselines. https://github.com/eth-lre/ verify-then-generate Teacher If the height is 6, what is the length of the box? Volume of a box is height * width * length. Student Multi-turn dialog tutoring task Goal: Generate next teacher utterance. Not quite. Is the length you computed 2-times more than height? targeted and correct A. Error reason (baseline): Student made a careless mistake. B. Correctness verification: incorrect C. Stepwise verification: Step 2 -We set an equation 2 * length = 6 ... D. Error Description: length is used as a label instead of a variable representing the number.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f87888f-51c1-438e-a342-31c7b33f1d82Cited by top-tier papers8
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based AgentsTao Wu, Jingyuan Chen, Wang Lin, Mengze Li et al.ACL 2025 · 16 citations
- Simulated Students in Tutoring Dialogues: Substance or Illusion?Alexander Scarlatos, Jaewook Lee, Simon Woodhead, Andrew LanACL 2026 · 6 citations
- MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM TutorsJakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur et al.EMNLP 2025 · 4 citations
- Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student AttacksJin Zhao, Marta Knezevic, Tanja KäserACL 2026 · 2 citations
- From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement LearningDavid Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi et al.EMNLP 2025
Builds on12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- Automatic Generation of Socratic Subquestions for Teaching Math Word ProblemsKumar Shridhar, Jakub Macina, Mennatallah El-Assady, Tanmay Sinha et al.EMNLP 2022 · 31 citations
Related papers
- Planning-Guided Tutoring with Assessment-Driven Memory for Pedagogical LLM TutorsZechen Li, Qiannan Zhu, Mei Wang, Jia Li et al.ACL 2026
- Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise GenerationZhenwen Liang, Wenhao Yu, Tanmay Rajpurohit, Peter Clark et al.EMNLP 2023 · 27 citations
- No Need for Explanations: LLMs can implicitly learn from mistakes in-contextLisa Alazraki, Maximilian Mozes, Jon Ander Campos, Yi Chern Tan et al.EMNLP 2025
- S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical ReasonersYuchen Yan, Jin Jiang, Yang Liu, Yixin Cao et al.AAAI 2025 · 19 citations
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 10 citations
