EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
Bin Xu, Yu Bai, Huashan Sun, Yiguan Lin, Siming Liu, Xinyue Liang, Yaolin Li, Zhuangzhi Dong, Jingren Zhang, Yufan Deng, Xinyu Zou, Yang Gao, Heyan Huang
摘要
As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational scenarios, incorporating synthetic data containing 9 major scenarios and over 4,000 distinct educational contexts. To enable comprehensive assessment, we propose a set of multi-dimensional evaluation rubrics that cover 12 critical aspects relevant to both teachers and students. We further apply human annotation to ensure the effectiveness of the model-generated evaluation responses. Additionally, we succeed to train a relatively small-scale model on our constructed dataset and demonstrate that it can achieve performance comparable to state-ofthe-art large models (e.g., Deepseek V3, Qwen Max) on the test set. Overall, this work provides a practical foundation for the development and evaluation of education-oriented language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated TeachersYilin Jiang, Mingzi Zhang, Xuanyu Yin, Sheng Jin 等AAAI 2026 · 被引用 1 次
- LongTutor: Benchmarking Large Language Models for Long-term Personalized TutoringNing Li, Zheng Zhang, Zhenya Huang, Rui Li 等ACL 2026
它引用的顶会 Paper3
- Domain-Adaptive Neural Automated Essay ScoringYue Cao, Hanqi Jin, Xiaojun Wan, Zhiwei YuSIGIR 2020 · 被引用 47 次
- EQG-RACE: Examination-Type Question GenerationXin Jia, Wenjie Zhou, Xu Sun, Yunfang WuAAAI 2021 · 被引用 43 次
- EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue FrameworkYao Shi, Rongkeng Liang, Yong XuACL 2025 · 被引用 18 次
相关 Paper
- K-12EduBench: A Benchmark for Evaluating Large Language Models' Knowledge, Problem-Solving, and Educational Goal Cognition in K-12 EducationYuqing Ye, Xuan Zhou, Zhifu Chen, Dandan Li 等AAAI 2026
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme 等ACL 2024 · 被引用 27 次
- RubricBench: Aligning Model-Generated Rubrics with Human StandardsJunyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu 等ACL 2026 · 被引用 7 次
- CFBench: A Comprehensive Constraints-Following Benchmark for LLMsTao Zhang, Chenglin Zhu, Yanjun Shen, Wenjing Luo 等ACL 2025 · 被引用 53 次
- CASE-Bench: Context-Aware SafEty Benchmark for Large Language ModelsGuangzhi Sun, Xiao Zhan, Shutong Feng, Philip C. Woodland 等ICML 2025
