From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
Weikang Shi, Houxing Ren, Junting Pan, Aojun Zhou, Ke Wang, Zimu Lu, Yunqiao Yang, Yuxuan Hu, Linda Wei, Mingjie Zhan, Hongsheng Li
摘要
Large Language Models (LLMs) show significant potential in AI mathematical tutoring, yet current evaluations often rely on simplistic metrics or narrow pedagogical scenarios, failing to assess comprehensive, multi-turn teaching effectiveness. In this paper, we introduce KMP-Bench, a comprehensive K-8 Mathematical Pedagogical Benchmark designed to assess LLMs from two complementary perspectives. The first module, KMP-Dialogue, evaluates holistic pedagogical capabilities against six core principles (e.g., Challenge, Explanation, Feedback), leveraging a novel multi-turn dialogue dataset constructed by weaving together diverse pedagogical components. The second module, KMP-Skills, provides a granular assessment of foundational tutoring abilities, including multi-turn problem-solving, error detection and correction, and problem generation. Our evaluations on KMP-Bench reveal a key disparity: while leading LLMs excel at tasks with verifiable solutions, they struggle with the nuanced application of pedagogical principles. Additionally, we present KMP-Pile, a large-scale (150K) dialogue dataset. Models fine-tuned on KMP-Pile show substantial improvement on KMP-Bench, underscoring the value of pedagogically-rich training data for developing more effective AI math tutors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- MMTutorBench: The First Multimodal Benchmark for AI Math TutoringTengchao Yang, Sichen Guo, Mengzhao Jia, Jiaming Su 等ACL 2026 · 被引用 2 次
- MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM TutorsJakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur 等EMNLP 2025 · 被引用 4 次
- LongTutor: Benchmarking Large Language Models for Long-term Personalized TutoringNing Li, Zheng Zhang, Zhenya Huang, Rui Li 等ACL 2026
- Simulated Students in Tutoring Dialogues: Substance or Illusion?Alexander Scarlatos, Jaewook Lee, Simon Woodhead, Andrew LanACL 2026 · 被引用 6 次
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He 等ACL 2024 · 被引用 35 次
