Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
Yuyan Chen, Songzhou Yan, Panjun Liu, Yanghua Xiao
摘要
Teachers are important to imparting knowledge and guiding learners, and the role of large language models (LLMs) as potential educators is emerging as an important area of study. Recognizing LLMs' capability to generate educational content can lead to advances in automated and personalized learning. While LLMs have been tested for their comprehension and problem-solving skills, their capability in teaching remains largely unexplored. In teaching, questioning is a key skill that guides students to analyze, evaluate, and synthesize core concepts and principles. Therefore, our research introduces a benchmark to evaluate the questioning capability in education as a teacher of LLMs through evaluating their generated educational questions, utilizing Anderson and Krathwohl's taxonomy across general, monodisciplinary, and interdisciplinary domains. We shift the focus from LLMs as learners to LLMs as educators, assessing their teaching capability through guiding them to generate questions. We apply four metrics, including relevance, coverage, representativeness, and consistency, to evaluate the educational quality of LLMs' outputs. Our results indicate that GPT-4 demonstrates significant potential in teaching general, humanities, and science courses; Claude2 appears more apt as an interdisciplinary teacher. Furthermore, the automatic scores align with human perspectives.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue FrameworkYao Shi, Rongkeng Liang, Yong XuACL 2025 · 被引用 18 次
- XMeCap: Meme Caption Generation with Sub-Image AdaptabilityYuyan Chen, Songzhou Yan, Zhihong Zhu, Zhixu Li 等ACM MM 2024 · 被引用 7 次
- SMART: Evaluating LLMs' Mathematical Reasoning via a Human Cognitive Process-Inspired BenchmarkYujie Hou, Mei Wang, Yaoyao Zhong, Ting Zhang 等ACL 2026
- SDBench: A Survey-based Domain-specific LLM Benchmarking and Optimization FrameworkCheng Guo, Hu Kai, Shuxian Liang, Yiyang Jiang 等ACL 2025
- K-12EduBench: A Benchmark for Evaluating Large Language Models' Knowledge, Problem-Solving, and Educational Goal Cognition in K-12 EducationYuqing Ye, Xuan Zhou, Zhifu Chen, Dandan Li 等AAAI 2026
它引用的顶会 Paper11
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu 等ICML 2024 · 被引用 220 次
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question GenerationYu Chen, Lingfei Wu, Mohammed J. ZakiICLR 2020 · 被引用 167 次
- Adapting Large Language Models via Reading ComprehensionDaixuan Cheng, Shaohan Huang, Furu WeiICLR 2024 · 被引用 146 次
相关 Paper
- Can Large Language Models Be Good Language Teachers?LiQing Xu, Qiwei Li, Tianshuo Peng, Zuchao Li 等EMNLP 2025
- SocraticLM: Exploring Socratic Personalized Teaching with Large Language ModelsJiayu Liu, Zhenya Huang, Tong Xiao, Jing Sha 等NeurIPS 2024 · 被引用 65 次
- CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language ModelsTong Zhang, Peixin Qin, Yang Deng, Chen Huang 等ACL 2024
- Simulated Students in Tutoring Dialogues: Substance or Illusion?Alexander Scarlatos, Jaewook Lee, Simon Woodhead, Andrew LanACL 2026 · 被引用 6 次
- I Could've Asked That: Reformulating Unanswerable QuestionsWenting Zhao, Ge Gao, Claire Cardie, Alexander M. RushEMNLP 2024 · 被引用 2 次
