Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise Generation
Zhenwen Liang, Wenhao Yu, Tanmay Rajpurohit, Peter Clark, Xiangliang Zhang, Ashwin Kalyan
摘要
In this paper, we present a novel approach for distilling math word problem solving capabilities from large language models (LLMs) into smaller, more efficient student models. Our approach is designed to consider the student model's weaknesses and foster a tailored learning experience by generating targeted exercises aligned with educational science principles, such as knowledge tracing and personalized learning. Concretely, we let GPT-3 be a math tutor and run two steps iteratively: 1) assessing the student model's current learning status on a GPT-generated exercise book, and 2) improving the student model by training it with tailored exercise samples generated by GPT-3. Experimental results reveal that our approach outperforms LLMs (e.g., GPT-3 and PaLM) in accuracy across three distinct benchmarks while employing significantly fewer parameters. Furthermore, we provide a comprehensive analysis of the various components within our methodology to substantiate their efficacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Cooperative Multi-Agent Framework for Zero-Shot Named Entity RecognitionZihan Wang, Ziqi Zhao, Yougang Lyu, Zhumin Chen 等WWW 2025 · 被引用 16 次
- Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral SimulationSonglin Xu, Hao-Ning Wen, Hongyi Pan, Dallas Dominguez 等CHI 2025 · 被引用 13 次
- Error-driven Data-efficient Large Multimodal Model TuningBarry Menglong Yao, Qifan Wang, Lifu HuangACL 2025 · 被引用 1 次
- SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool CallingFares Fawzi, Vinitra Swamy, Dominik Glandorf, Tanya Nazaretsky 等EMNLP 2025
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei 等ICLR 2023 · 被引用 318 次
相关 Paper
- MathScale: Scaling Instruction Tuning for Mathematical ReasoningZhengyang Tang, Xingxing Zhang, Benyou Wang, Furu WeiICML 2024 · 被引用 163 次
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 被引用 10 次
- Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model TutorsNico Daheim, Jakub Macina, Manu Kapur, Iryna Gurevych 等EMNLP 2024 · 被引用 1 次
- Template-Theorems Graph Construction to Enhance Mathematical Reasoning Capabilities of LLMYarong Lan, Yajing Xu, Huajun ChenAAAI 2026
- MMTutorBench: The First Multimodal Benchmark for AI Math TutoringTengchao Yang, Sichen Guo, Mengzhao Jia, Jiaming Su 等ACL 2026 · 被引用 2 次
