Can Large Language Models Be Good Language Teachers?
LiQing Xu, Qiwei Li, Tianshuo Peng, Zuchao Li, Hai Zhao, Ping Wang
摘要
Large language models (LLMs) have achieved remarkable success across diverse domains. However, their potential as effective language teachers-particularly in complex pedagogical scenarios like teaching Chinese as a second language-remains inadequately assessed. To address this gap, we propose the first pedagogical competence benchmark for LLMs, rigorously evaluating their performance against international standards for Chinese language teachers. Our framework spans three core dimensions: (1) basic knowledge evaluation, covering 32 subtopics across five major categories; (2) international teacher examination, based on data collected from international Chinese teacher certification exams; and (3) teaching practice evaluation, where target LLMs summarize knowledge points and design instructional content for student models, followed by testing the student models to assess the LLM's ability to distill and teach key concepts. We conduct a comprehensive evaluation of 13 latest multilingual and Chinese LLMs. While most models demonstrate promising pedagogical potential, there remains substantial room for improvement in their teaching capabilities. This study contributes to the development of AI-assisted language education tools capable of rivaling human teaching excellence. The benchmark dataset and evaluation scripts used in this study are publicly available at https: //github.com/Line-Kite/CLTE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
相关 Paper
- MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language ModelsYan Cai, Linlin Wang, Ye Wang, Gerard de Melo 等AAAI 2024 · 被引用 42 次
- CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical ScenariosZetian Ouyang, Yishuai Qiu, Linlin Wang, Gerard de Melo 等EMNLP 2024 · 被引用 6 次
- AlignBench: Benchmarking Chinese Alignment of Large Language ModelsXiao Liu, Xuanyu Lei, Shengyuan Wang, Yue Huang 等ACL 2024 · 被引用 9 次
- CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science MasteryXiaoshuai Song, Muxi Diao, Guanting Dong, Zhengyang Wang 等ICLR 2025
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun 等ACL 2024
