Can Large Language Models Be Good Language Teachers?
LiQing Xu, Qiwei Li, Tianshuo Peng, Zuchao Li, Hai Zhao, Ping Wang
Abstract
Large language models (LLMs) have achieved remarkable success across diverse domains. However, their potential as effective language teachers-particularly in complex pedagogical scenarios like teaching Chinese as a second language-remains inadequately assessed. To address this gap, we propose the first pedagogical competence benchmark for LLMs, rigorously evaluating their performance against international standards for Chinese language teachers. Our framework spans three core dimensions: (1) basic knowledge evaluation, covering 32 subtopics across five major categories; (2) international teacher examination, based on data collected from international Chinese teacher certification exams; and (3) teaching practice evaluation, where target LLMs summarize knowledge points and design instructional content for student models, followed by testing the student models to assess the LLM's ability to distill and teach key concepts. We conduct a comprehensive evaluation of 13 latest multilingual and Chinese LLMs. While most models demonstrate promising pedagogical potential, there remains substantial room for improvement in their teaching capabilities. This study contributes to the development of AI-assisted language education tools capable of rivaling human teaching excellence. The benchmark dataset and evaluation scripts used in this study are publicly available at https: //github.com/Line-Kite/CLTE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e7580a0-b316-4150-9ed7-a0a32633a0cfBuilds on3
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
Related papers
- MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language ModelsYan Cai, Linlin Wang, Ye Wang, Gerard de Melo et al.AAAI 2024 · 42 citations
- CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical ScenariosZetian Ouyang, Yishuai Qiu, Linlin Wang, Gerard de Melo et al.EMNLP 2024 · 6 citations
- AlignBench: Benchmarking Chinese Alignment of Large Language ModelsXiao Liu, Xuanyu Lei, Shengyuan Wang, Yue Huang et al.ACL 2024 · 9 citations
- CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science MasteryXiaoshuai Song, Muxi Diao, Guanting Dong, Zhengyang Wang et al.ICLR 2025
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
