LongTutor: Benchmarking Large Language Models for Long-term Personalized Tutoring
Ning Li, Zheng Zhang, Zhenya Huang, Rui Li, Yi Zhan, Yinbo Luo, Qi Liu, Enhong Chen
摘要
The rapid advancement of large language models (LLMs) has driven the deployment of LLMbased AI tutors on online learning platforms. This widespread adoption highlights an urgent need for systematic benchmarks to evaluate their tutoring capabilities. However, existing evaluations predominantly focus on isolated, short-term interactions, overlooking the inherently long-term nature of learning. To bridge this gap, we introduce LongTutor, a benchmark for long-term personalized tutoring grounded in formative assessment theory. Built from expertannotated real-world learning logs, LongTutor evaluates LLMs across three progressive tasks: historical evidence acquisition, knowledge state diagnosis, and adaptive teaching action. Our experiments reveal a critical capability mismatch: while LLMs excel at evidence acquisition, they struggle to effectively leverage long-term history for accurate diagnosis and adaptive teaching. To enable scalable benchmark expansion, we further propose an automated generator-verifier pipeline, paving the way toward truly long-term AI tutoring systems. We release our code and dataset at https://github.com/liano3/LongTutor .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
相关 Paper
- MMTutorBench: The First Multimodal Benchmark for AI Math TutoringTengchao Yang, Sichen Guo, Mengzhao Jia, Jiaming Su 等ACL 2026 · 被引用 2 次
- From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-BenchWeikang Shi, Houxing Ren, Junting Pan, Aojun Zhou 等AAAI 2026
- Simulated Students in Tutoring Dialogues: Substance or Illusion?Alexander Scarlatos, Jaewook Lee, Simon Woodhead, Andrew LanACL 2026 · 被引用 6 次
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive MemoryDi Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang 等ICLR 2025
- A Theory of Adaptive Scaffolding for LLM-Based Pedagogical AgentsClayton Cohn, Surya Rayala, Namrata Srivastava, Joyce Horn Fonteles 等AAAI 2026 · 被引用 4 次
