CamemBERT: a Tasty French Language Model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, Benoît Sagot
Abstract
Pretrained language models are now ubiquitous in Natural Language Processing. Despite their success, most available models have either been trained on English data or on the concatenation of data in multiple languages. This makes practical use of such models-in all languages except English-very limited. In this paper, we investigate the feasibility of training monolingual Transformer-based language models for other languages, taking French as an example and evaluating our language models on part-of-speech tagging, dependency parsing, named entity recognition and natural language inference tasks. We show that the use of web crawled data is preferable to the use of Wikipedia data. More surprisingly, we show that a relatively small web crawled dataset (4GB) leads to results that are as good as those obtained using larger datasets (130+GB). Our best performing model CamemBERT reaches or improves the state of the art in all four downstream tasks. 1 Released at: https://camembert-model.fr under the MIT open-source license.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- CASSINI: Network-Aware Job Scheduling in Machine Learning ClustersSudarsanan Rajasekaran, Manya Ghobadi, Aditya AkellaNSDI 2024 · 144 citations
- MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transferIlias Chalkidis, Manos Fergadiotis, Ion AndroutsopoulosEMNLP 2021 · 78 citations
- A Statutory Article Retrieval Dataset in FrenchAntoine Louis, Gerasimos SpanakisACL 2022 · 59 citations
- MAGMA: An Optimization Framework for Mapping Multiple DNNs on Multiple Accelerator CoresSheng-Chun Kao, Tushar KrishnaHPCA 2022 · 58 citations
Builds on2
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domainsYanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier et al.ACL 2023 · 19 citations
- A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource LanguagesPedro Javier Ortiz Suárez, Laurent Romary, Benoît SagotACL 2020 · 72 citations
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 235 citations
- Improving Low-Resource Languages in Pre-Trained Multilingual Language ModelsViktor Hangya, Hossain Shaikh Saadi, Alexander FraserEMNLP 2022 · 17 citations
- Multilingual Language Model Pretraining using Machine-translated DataJiayi Wang, Yao Lu, Maurice Weber, Max Ryabinin et al.EMNLP 2025
