Simulated Students in Tutoring Dialogues: Substance or Illusion?
Alexander Scarlatos, Jaewook Lee, Simon Woodhead, Andrew Lan
Abstract
Advances in large language models (LLMs) enable many new innovations in education. However, evaluating the effectiveness of new technology requires real students, which is time-consuming and hard to scale up. Therefore, many recent works on LLM-powered tutoring solutions have used simulated students for both training and evaluation, often via simple prompting. Surprisingly, little work has been done to ensure or even measure the quality of simulated students. In this work, we formally define the student simulation task, propose a set of evaluation metrics that span linguistic, behavioral, and cognitive aspects, and benchmark a wide range of student simulation methods on these metrics. We experiment on a real-world math tutoring dialogue dataset, where both automated and human evaluation results show that prompting strategies for student simulation perform poorly; supervised fine-tuning and preference optimization yield much better but still limited performance, motivating future work on this challenging task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 518d2c69-cf2d-41d5-baad-02ea392a9695Builds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Teach AI How to Code: Using Large Language Models as Teachable Agents for Programming EducationHyoungwook Jin, Seonghee Lee, Hyungyu Shin, Juho KimCHI 2024 · 94 citations
- Consistently Simulating Human Personas with Multi-Turn Reinforcement LearningMarwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff et al.NeurIPS 2025 · 51 citations
Related papers
- Personality-aware Student Simulation for Conversational Intelligent Tutoring SystemsZhengyuan Liu, Stella Xin Yin, Geyu Lin, Nancy F. ChenEMNLP 2024 · 13 citations
- Planning-Guided Tutoring with Assessment-Driven Memory for Pedagogical LLM TutorsZechen Li, Qiannan Zhu, Mei Wang, Jia Li et al.ACL 2026
- MMTutorBench: The First Multimodal Benchmark for AI Math TutoringTengchao Yang, Sichen Guo, Mengzhao Jia, Jiaming Su et al.ACL 2026 · 2 citations
- SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie et al.EMNLP 2025
- TeachTune: Reviewing Pedagogical Agents Against Diverse Student Profiles with Simulated StudentsHyoungwook Jin, Minju Yoo, Jeongeon Park, Yokyung Lee et al.CHI 2025 · 58 citations
