MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors
Jakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan
Abstract
Evaluating the pedagogical capabilities of AIbased tutoring models is critical for making guided progress in the field. Yet, we lack a reliable, easy-to-use, and simple-to-run evaluation that reflects the pedagogical abilities of models. To fill this gap, we present MATH-TUTORBENCH, an open-source benchmark for holistic tutoring model evaluation. MATHTU-TORBENCH contains a collection of datasets and metrics that broadly cover tutor abilities as defined by learning sciences research in dialogbased teaching. To score the pedagogical quality of open-ended teacher responses, we train a reward model and show it can discriminate expert from novice teacher responses with high accuracy. We evaluate a wide set of closed-and open-weight models on MATHTUTORBENCH and find that subject expertise, indicated by solving ability, does not immediately translate to good teaching. Rather, pedagogy and subject expertise appear to form a trade-off that is navigated by the degree of tutoring specialization of the model. Furthermore, tutoring appears to become more challenging in longer dialogs, where simpler questioning strategies begin to fail. We release the benchmark, code, and leaderboard openly to enable rapid benchmarking of future models. 1 github.com/eth-lre/mathtutorbench
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15f5dac0-59e0-40af-862d-611a2f13ebf0Cited by top-tier papers5
- MMTutorBench: The First Multimodal Benchmark for AI Math TutoringTengchao Yang, Sichen Guo, Mengzhao Jia, Jiaming Su et al.ACL 2026 · 2 citations
- Position: LLMs Can be Good Tutors in English EducationJingheng Ye, Shen Wang, Deqing Zou, Yibo Yan et al.EMNLP 2025 · 2 citations
- From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement LearningDavid Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi et al.EMNLP 2025
- LongTutor: Benchmarking Large Language Models for Long-term Personalized TutoringNing Li, Zheng Zhang, Zhenya Huang, Rui Li et al.ACL 2026
- K-12EduBench: A Benchmark for Evaluating Large Language Models' Knowledge, Problem-Solving, and Educational Goal Cognition in K-12 EducationYuqing Ye, Xuan Zhou, Zhifu Chen, Dandan Li et al.AAAI 2026
Builds on9
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- Dialog Inpainting: Turning Documents into DialogsZhuyun Dai, Arun Tejasvi Chaganty, Vincent Y. Zhao, Aida Amini et al.ICML 2022 · 77 citations
- SocraticLM: Exploring Socratic Personalized Teaching with Large Language ModelsJiayu Liu, Zhenya Huang, Tong Xiao, Jing Sha et al.NeurIPS 2024 · 65 citations
Related papers
- From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-BenchWeikang Shi, Houxing Ren, Junting Pan, Aojun Zhou et al.AAAI 2026
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon et al.ICLR 2026 · 13 citations
- UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language ModelsXin Xu, Jiaxin Zhang, Tianhao Chen, Zitong Chao et al.ICLR 2025
- VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across DomainsXuzhao Li, Xuchen Li, Shiyu Hu, Yongzhen Guo et al.AAAI 2026 · 16 citations
- EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue FrameworkYao Shi, Rongkeng Liang, Yong XuACL 2025 · 18 citations
