TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay Scoring
Rodrigo Wilkens, Alice Pintard, David Alfter, Vincent Folny, Thomas François
Abstract
Automated Essay Scoring (AES) aims to automatically assess the quality of essays. Automation enables large-scale assessment, improvements in consistency, reliability, and standardization. Those characteristics are of particular relevance in the context of language certification exams. However, a major bottleneck in the development of AES systems is the availability of corpora, which, unfortunately, are scarce, especially for languages other than English. In this paper, we aim to foster the development of AES for French by providing the TCFLE-8 corpus, a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF -French Knowledge Test) certification exam. We report the strict quality procedure that led to the scoring of each essay by at least two raters according to the levels of the Common European Framework of Reference for Languages (CEFR) and to the creation of a balanced corpus. In addition, we describe how linguistic properties of the essays relate to the learners' proficiency in TCFLE-8. We also advance the state-of-the-art performance for the AES task in French by experimenting with two strong baselines (i.e., RoBERTa and featurebased). Finally, we discuss the challenges of AES using TCFLE-8. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab2e1959-7287-49db-81da-e6e74792cee8Cited by top-tier papers2
- DREsS: Dataset for Rubric-based Essay Scoring on EFL WritingHaneul Yoo, Jieun Han, So-Yeon Ahn, Alice OhACL 2025 · 13 citations
- UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency AssessmentJoseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens et al.EMNLP 2025 · 2 citations
Builds on1
Related papers
- Conundrums in Cross-Prompt Automated Essay Scoring: Making Sense of the State of the ArtShengjie Li, Vincent NgACL 2024 · 8 citations
- Cross-Prompt Automated Essay Scoring of Multiple Traits: Making Sense of the State of the ArtShengjie Li, Vincent NgACL 2026 · 10 citations
- CEFR-Based Sentence Difficulty Annotation and AssessmentYuki Arase, Satoru Uchida, Tomoyuki KajiwaraEMNLP 2022 · 17 citations
- Improving Domain Generalization for Prompt-Aware Essay Scoring via Disentangled Representation LearningZhiwei Jiang, Tianyi Gao, Yafeng Yin, Meng Liu et al.ACL 2023 · 16 citations
- Mixture of Ordered Scoring Experts for Cross-prompt Essay Trait ScoringPo-Kai Chen, Bo-Wei Tsai, Shao-Kuan Wei, Chien-Yao Wang et al.ACL 2025
