QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
David Beauchemin, Richard Khoury
Abstract
Large language models (LLM) perform outstandingly in various downstream tasks. However, there is limited understanding regarding how these models internalize linguistic knowledge, so various linguistic benchmarks have recently been proposed to facilitate syntactic evaluation of language models (LM) across languages. This paper introduces QFrCoLA (Quebec-French Corpus of Linguistic Acceptability Judgments), a normative binary acceptability judgments dataset comprising 25,153 in-domain and 2,675 out-of-domain sentences. Our study leverages the QFrCoLA dataset and seven other linguistic binary acceptability judgments corpus to benchmark eight LM. The results demonstrate that, on average, finetuned Transformer-based LM are strong baselines for most languages and that zero-shot binary classification LLM perform worse than the naive baseline on the task. However, for the QFrCoLA benchmark, on average, a finetuned Transformer-based LM outperformed other methods tested. It also shows that pretrained cross-lingual LLMs selected for our experimentation do not seem to have acquired linguistic judgment capabilities during their pre-training for Quebec French. Finally, our experiment results on QFrCoLA show that our dataset, built from examples that illustrate linguistic norms rather than speakers' feelings, is similar to linguistic acceptability judgment; it is a challenging dataset that can benchmark LM on their linguistic judgment capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- BLOOM+1: Adding Language Support to BLOOM for Zero-Shot PromptingZheng Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji et al.ACL 2023 · 20 citations
- RuCoLA: Russian Corpus of Linguistic AcceptabilityVladislav Mikhailov, Tatiana Shamardina, Max Ryabinin, Alena Pestova et al.EMNLP 2022 · 19 citations
- BLESS: Benchmarking Large Language Models on Sentence SimplificationTannon Kew, Alison Chi, Laura Vásquez-Rodríguez, Sweta Agrawal et al.EMNLP 2023 · 15 citations
Related papers
- MELA: Multilingual Evaluation of Linguistic AcceptabilityZiyin Zhang, Yikang Liu, Weifang Huang, Junyu Mao et al.ACL 2024
- F-Eval: Asssessing Fundamental Abilities with Refined Evaluation MethodsYu Sun, Keyuchen Keyuchen, Shujie Wang, Peiji Li et al.ACL 2024
- MedAraBench: Large-scale Arabic Medical Question Answering Dataset and BenchmarkMouath Abu Daoud, Leen Kharouf, Omar El Hajj, Dana El Samad et al.ICLR 2026 · 4 citations
- KoLA: Carefully Benchmarking World Knowledge of Large Language ModelsJifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao et al.ICLR 2024 · 91 citations
- Truth Knows No Language: Evaluating Truthfulness Beyond EnglishBlanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes et al.ACL 2025
