SLING: Sino Linguistic Evaluation of Large Language Models
Yixiao Song, Kalpesh Krishna, Rajesh Bhatt, Mohit Iyyer
Abstract
To understand what kinds of linguistic knowledge are encoded by pretrained Chinese language models (LMs), we introduce the benchmark of Sino LINGuistics (SLING), which consists of 38K minimal sentence pairs in Mandarin Chinese grouped into 9 high-level linguistic phenomena. Each pair demonstrates the acceptability contrast of a specific syntactic or semantic phenomenon (e.g., The keys are lost vs. The keys is lost), and an LM should assign lower perplexity to the acceptable sentence. In contrast to the CLiMP dataset (Xiang et al., 2021), which also contains Chinese minimal pairs and was created by translating the vocabulary of the English BLiMP dataset, the minimal pairs in SLING are derived primarily by applying syntactic and lexical transformations to naturallyoccurring, linguist-annotated sentences from the Chinese Treebank 9.0, thus addressing severe issues in CLiMP's data generation process. We test 18 publicly available pretrained monolingual (e.g., BERT-base-zh, CPM) and multi-lingual (e.g., mT5, XLM) language models on SLING. Our experiments show that the average accuracy for LMs is far below human performance (69.7% vs. 97.1%), while BERT-base-zh achieves the highest accuracy (84.8%) of all tested LMs, even much larger ones. Additionally, we find that most LMs have a strong gender and number (singular/plural) bias, and they perform better on local phenomena than hierarchical ones. 1 A: 他们在吃饭了。 (They are already in the process of having a meal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Minimal Pair-Based Evaluation of Code-SwitchingIgor Sterner, Simone TeufelACL 2025 · 8 citations
- Different types of syntactic agreement recruit the same units within large language modelsDaria Kryvosheieva, Andrea Gregor de Varda, Evelina Fedorenko, Greta TuckuteACL 2026 · 3 citations
- RuBLiMP: Russian Benchmark of Linguistic Minimal PairsEkaterina Taktasheva, Maxim Bazhukov, Kirill Koncha, Alena Fenogenova et al.EMNLP 2024 · 2 citations
- MELA: Multilingual Evaluation of Linguistic AcceptabilityZiyin Zhang, Yikang Liu, Weifang Huang, Junyu Mao et al.ACL 2024
- Implicit Representations of Grammaticality in Language ModelsYingshan Susan Wang, Linlu Qiu, Zhaofeng Wu, Roger P. Levy et al.ACL 2026
Builds on4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 34 citations
- Surface Form Competition: Why the Highest Probability Answer Isn't Always RightAri Holtzman, Peter West, Vered Shwartz, Yejin Choi et al.EMNLP 2021 · 10 citations
Related papers
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability JudgmentsDavid Beauchemin, Richard KhouryEMNLP 2025
- Enhancing Chinese Pre-trained Language Model via Heterogeneous Linguistics GraphYanzeng Li, Jiangxia Cao, Xin Cong, Zhenyu Zhang et al.ACL 2022 · 11 citations
- CHENGYU-BENCH: Benchmarking Large Language Models for Chinese Idiom Understanding and UseYicheng Fu, Zhemin Huang, Liuxin Yang, Yumeng Lu et al.EMNLP 2025
- TurBLiMP: A Turkish Benchmark of Linguistic Minimal PairsEzgi Basar, Francesca Padovani, Jaap Jumelet, Arianna BisazzaEMNLP 2025
- XSemPLR: Cross-Lingual Semantic Parsing in Multiple Natural Languages and Meaning RepresentationsYusen Zhang, Jun Wang, Zhiguo Wang, Rui ZhangACL 2023 · 4 citations
