K-12EduBench: A Benchmark for Evaluating Large Language Models' Knowledge, Problem-Solving, and Educational Goal Cognition in K-12 Education
Yuqing Ye, Xuan Zhou, Zhifu Chen, Dandan Li, Hengnian Gu, Jin Peng Zhou, Dongdai Zhou
Abstract
Large language models hold great promise for transforming K-12 education, but there is an urgent need for systematic evaluation of their core educational capabilities. Existing benchmarks often overlook educational goal cognition and overemphasize answer accuracy, thereby failing to capture deeper subject-level knowledge ability and problem-solving ability. To address this gap, we introduce K-12EduBench: a benchmark for evaluating LLMs’ subject-level knowledge ability, subject-specific problem-solving ability, and educational goal cognition ability in K-12 education. K-12EduBench comprises four components: (1) a dataset of 2,640 objective and 619 subjective questions across nine subjects, annotated with answers, problem-solving processes, and cognitive-level labels; (2) nine Item Response Theory (IRT) models for estimating subject-level knowledge ability; (3) evaluation methods and metrics for assessing multi-step problem-solving ability; and (4) prompts and scoring rubrics for measuring alignment with target cognitive levels. Experiments on advanced LLMs show that education-optimized models consistently outperform general-purpose ones across all three abilities, while under-scaled models lag substantially. We observe a strong positive correlation between subject-level knowledge ability and subject-specific problem-solving ability. Despite gains in educational goal cognition ability, current models—even those tailored for education—still fall short of real-world instructional needs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf56356f-071f-489a-9df9-be451b43b67bBuilds on6
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
- Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge EvaluationZhouhong Gu, Xiaoxuan Zhu, Haoning Ye, Lin Zhang et al.AAAI 2024 · 82 citations
- Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language ModelsYuyan Chen, Songzhou Yan, Panjun Liu, Yanghua XiaoACL 2024 · 8 citations
- MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM TutorsJakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur et al.EMNLP 2025 · 4 citations
Related papers
- EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMsNumaan Naeem, Abdellah El Mekki, Muhammad Abdul-MageedEMNLP 2025
- CK12: A Rounded K12 Knowledge Graph Based Benchmark for Chinese Holistic Cognition EvaluationWeihao You, Pengcheng Wang, Changlong Li, Zhilong Ji et al.AAAI 2024 · 4 citations
- From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-BenchWeikang Shi, Houxing Ren, Junting Pan, Aojun Zhou et al.AAAI 2026
- UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language ModelsXin Xu, Jiaxin Zhang, Tianhao Chen, Zitong Chao et al.ICLR 2025
- EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational ScenariosBin Xu, Yu Bai, Huashan Sun, Yiguan Lin et al.ACL 2026
