EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
Bin Xu, Yu Bai, Huashan Sun, Yiguan Lin, Siming Liu, Xinyue Liang, Yaolin Li, Zhuangzhi Dong, Jingren Zhang, Yufan Deng, Xinyu Zou, Yang Gao, Heyan Huang
Abstract
As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational scenarios, incorporating synthetic data containing 9 major scenarios and over 4,000 distinct educational contexts. To enable comprehensive assessment, we propose a set of multi-dimensional evaluation rubrics that cover 12 critical aspects relevant to both teachers and students. We further apply human annotation to ensure the effectiveness of the model-generated evaluation responses. Additionally, we succeed to train a relatively small-scale model on our constructed dataset and demonstrate that it can achieve performance comparable to state-ofthe-art large models (e.g., Deepseek V3, Qwen Max) on the test set. Overall, this work provides a practical foundation for the development and evaluation of education-oriented language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 546d8d57-32b1-4d16-bcc6-022e4d90ead4Cited by top-tier papers2
- EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated TeachersYilin Jiang, Mingzi Zhang, Xuanyu Yin, Sheng Jin et al.AAAI 2026 · 1 citation
- LongTutor: Benchmarking Large Language Models for Long-term Personalized TutoringNing Li, Zheng Zhang, Zhenya Huang, Rui Li et al.ACL 2026
Builds on3
- Domain-Adaptive Neural Automated Essay ScoringYue Cao, Hanqi Jin, Xiaojun Wan, Zhiwei YuSIGIR 2020 · 47 citations
- EQG-RACE: Examination-Type Question GenerationXin Jia, Wenjie Zhou, Xu Sun, Yunfang WuAAAI 2021 · 43 citations
- EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue FrameworkYao Shi, Rongkeng Liang, Yong XuACL 2025 · 18 citations
Related papers
- K-12EduBench: A Benchmark for Evaluating Large Language Models' Knowledge, Problem-Solving, and Educational Goal Cognition in K-12 EducationYuqing Ye, Xuan Zhou, Zhifu Chen, Dandan Li et al.AAAI 2026
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme et al.ACL 2024 · 27 citations
- RubricBench: Aligning Model-Generated Rubrics with Human StandardsJunyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu et al.ACL 2026 · 7 citations
- CFBench: A Comprehensive Constraints-Following Benchmark for LLMsTao Zhang, Chenglin Zhu, Yanjun Shen, Wenjing Luo et al.ACL 2025 · 53 citations
- CASE-Bench: Context-Aware SafEty Benchmark for Large Language ModelsGuangzhi Sun, Xiao Zhan, Shutong Feng, Philip C. Woodland et al.ICML 2025
