SLING: Sino Linguistic Evaluation of Large Language Models
Yixiao Song, Kalpesh Krishna, Rajesh Bhatt, Mohit Iyyer
摘要
To understand what kinds of linguistic knowledge are encoded by pretrained Chinese language models (LMs), we introduce the benchmark of Sino LINGuistics (SLING), which consists of 38K minimal sentence pairs in Mandarin Chinese grouped into 9 high-level linguistic phenomena. Each pair demonstrates the acceptability contrast of a specific syntactic or semantic phenomenon (e.g., The keys are lost vs. The keys is lost), and an LM should assign lower perplexity to the acceptable sentence. In contrast to the CLiMP dataset (Xiang et al., 2021), which also contains Chinese minimal pairs and was created by translating the vocabulary of the English BLiMP dataset, the minimal pairs in SLING are derived primarily by applying syntactic and lexical transformations to naturallyoccurring, linguist-annotated sentences from the Chinese Treebank 9.0, thus addressing severe issues in CLiMP's data generation process. We test 18 publicly available pretrained monolingual (e.g., BERT-base-zh, CPM) and multi-lingual (e.g., mT5, XLM) language models on SLING. Our experiments show that the average accuracy for LMs is far below human performance (69.7% vs. 97.1%), while BERT-base-zh achieves the highest accuracy (84.8%) of all tested LMs, even much larger ones. Additionally, we find that most LMs have a strong gender and number (singular/plural) bias, and they perform better on local phenomena than hierarchical ones. 1 A: 他们在吃饭了。 (They are already in the process of having a meal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Minimal Pair-Based Evaluation of Code-SwitchingIgor Sterner, Simone TeufelACL 2025 · 被引用 8 次
- Different types of syntactic agreement recruit the same units within large language modelsDaria Kryvosheieva, Andrea Gregor de Varda, Evelina Fedorenko, Greta TuckuteACL 2026 · 被引用 3 次
- RuBLiMP: Russian Benchmark of Linguistic Minimal PairsEkaterina Taktasheva, Maxim Bazhukov, Kirill Koncha, Alena Fenogenova 等EMNLP 2024 · 被引用 2 次
- MELA: Multilingual Evaluation of Linguistic AcceptabilityZiyin Zhang, Yikang Liu, Weifang Huang, Junyu Mao 等ACL 2024
- Implicit Representations of Grammaticality in Language ModelsYingshan Susan Wang, Linlu Qiu, Zhaofeng Wu, Roger P. Levy 等ACL 2026
它引用的顶会 Paper4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 被引用 34 次
- Surface Form Competition: Why the Highest Probability Answer Isn't Always RightAri Holtzman, Peter West, Vered Shwartz, Yejin Choi 等EMNLP 2021 · 被引用 10 次
相关 Paper
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability JudgmentsDavid Beauchemin, Richard KhouryEMNLP 2025
- Enhancing Chinese Pre-trained Language Model via Heterogeneous Linguistics GraphYanzeng Li, Jiangxia Cao, Xin Cong, Zhenyu Zhang 等ACL 2022 · 被引用 11 次
- CHENGYU-BENCH: Benchmarking Large Language Models for Chinese Idiom Understanding and UseYicheng Fu, Zhemin Huang, Liuxin Yang, Yumeng Lu 等EMNLP 2025
- TurBLiMP: A Turkish Benchmark of Linguistic Minimal PairsEzgi Basar, Francesca Padovani, Jaap Jumelet, Arianna BisazzaEMNLP 2025
- XSemPLR: Cross-Lingual Semantic Parsing in Multiple Natural Languages and Meaning RepresentationsYusen Zhang, Jun Wang, Zhiguo Wang, Rui ZhangACL 2023 · 被引用 4 次
