Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean
SungHo Kim, Nayeon Kim, Taehee Jeon, SangKeun Lee
Abstract
We introduce the Korean Grammar Evaluation BenchMark (KoGEM), designed to assess the linguistic competence of LLMs and humans in Korean. KoGEM consists of 1.5k multiplechoice QA pairs covering five main categories and 16 subcategories. The zero-shot evaluation of 27 LLMs of various sizes and types reveals that while LLMs perform remarkably well on straightforward tasks requiring primarily definitional knowledge, they struggle with tasks that demand the integration of realworld experiential knowledge, such as phonological rules and pronunciation. Furthermore, our in-depth analysis suggests that incorporating such experiential knowledge could enhance the linguistic competence of LLMs. With Ko-GEM, we not only highlight the limitations of current LLMs in linguistic competence but also uncover hidden facets of LLMs in linguistic competence, paving the way for enhancing comprehensive language understanding. Our code and dataset are available at: https://github.com/SungHo3268/KoGEM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17227640-6213-4c20-9299-6bcdf2d6d73bCited by top-tier papers1
Ask how each one uses itBuilds on9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- The Geometry of Multilingual Language Model RepresentationsTyler A. Chang, Zhuowen Tu, Benjamin K. BergenEMNLP 2022 · 22 citations
- Information-Theoretic Probing for Linguistic StructureTiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod et al.ACL 2020 · 21 citations
- ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language ModelsAparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati et al.ACL 2024 · 19 citations
Related papers
- TASE: Token Awareness and Structured Evaluation for Multilingual Language ModelsChenzhuo Zhao, Xinda Wang, Yue Huang, Junting Lu et al.AAAI 2026 · 1 citation
- KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language ModelsYuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan et al.WWW 2024 · 6 citations
- CK12: A Rounded K12 Knowledge Graph Based Benchmark for Chinese Holistic Cognition EvaluationWeihao You, Pengcheng Wang, Changlong Li, Zhilong Ji et al.AAAI 2024 · 4 citations
- Can Large Language Models Be Good Language Teachers?LiQing Xu, Qiwei Li, Tianshuo Peng, Zuchao Li et al.EMNLP 2025
- CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science MasteryXiaoshuai Song, Muxi Diao, Guanting Dong, Zhengyang Wang et al.ICLR 2025
