Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
Nico Daheim, Jakub Macina, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan
摘要
Large language models (LLMs) present an opportunity to scale high-quality personalized education to all. A promising approach towards this means is to build dialog tutoring models that scaffold students' problem-solving. However, even though existing LLMs perform well in solving reasoning questions, they struggle to precisely detect student's errors and tailor their feedback to these errors. Inspired by realworld teaching practice where teachers identify student errors and customize their response based on them, we focus on verifying student solutions and show how grounding to such verification improves the overall quality of tutor response generation. We collect a dataset of 1K stepwise math reasoning chains with the first error step annotated by teachers. We show empirically that finding the mistake in a student solution is challenging for current models. We propose and evaluate several verifiers for detecting these errors. Using both automatic and human evaluation we show that the student solution verifiers steer the generation model towards highly targeted responses to student errors which are more often correct with less hallucinations compared to existing baselines. https://github.com/eth-lre/ verify-then-generate Teacher If the height is 6, what is the length of the box? Volume of a box is height * width * length. Student Multi-turn dialog tutoring task Goal: Generate next teacher utterance. Not quite. Is the length you computed 2-times more than height? targeted and correct A. Error reason (baseline): Student made a careless mistake. B. Correctness verification: incorrect C. Stepwise verification: Step 2 -We set an equation 2 * length = 6 ... D. Error Description: length is used as a label instead of a variable representing the number.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based AgentsTao Wu, Jingyuan Chen, Wang Lin, Mengze Li 等ACL 2025 · 被引用 16 次
- Simulated Students in Tutoring Dialogues: Substance or Illusion?Alexander Scarlatos, Jaewook Lee, Simon Woodhead, Andrew LanACL 2026 · 被引用 6 次
- MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM TutorsJakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur 等EMNLP 2025 · 被引用 4 次
- Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student AttacksJin Zhao, Marta Knezevic, Tanja KäserACL 2026 · 被引用 2 次
- From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement LearningDavid Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi 等EMNLP 2025
它引用的顶会 Paper12
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
- Automatic Generation of Socratic Subquestions for Teaching Math Word ProblemsKumar Shridhar, Jakub Macina, Mennatallah El-Assady, Tanmay Sinha 等EMNLP 2022 · 被引用 31 次
相关 Paper
- Planning-Guided Tutoring with Assessment-Driven Memory for Pedagogical LLM TutorsZechen Li, Qiannan Zhu, Mei Wang, Jia Li 等ACL 2026
- Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise GenerationZhenwen Liang, Wenhao Yu, Tanmay Rajpurohit, Peter Clark 等EMNLP 2023 · 被引用 27 次
- No Need for Explanations: LLMs can implicitly learn from mistakes in-contextLisa Alazraki, Maximilian Mozes, Jon Ander Campos, Yi Chern Tan 等EMNLP 2025
- S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical ReasonersYuchen Yan, Jin Jiang, Yang Liu, Yixin Cao 等AAAI 2025 · 被引用 19 次
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 被引用 10 次
