LeXFiles and LegalLAMA: Facilitating English Multinational Legal Language Model Development
Ilias Chalkidis, Nicolas Garneau, Catalina Goanta, Daniel Martin Katz, Anders Søgaard
Abstract
In this work, we conduct a detailed analysis on the performance of legal-oriented pre-trained language models (PLMs). We examine the interplay between their original objective, acquired knowledge, and legal language understanding capacities which we define as the upstream, probing, and downstream performance, respectively. We consider not only the models' size but also the pre-training corpora used as important dimensions in our study. To this end, we release a multinational English legal corpus (LeXFiles) and a legal knowledge probing benchmark (LegalLAMA) to facilitate training and detailed analysis of legal-oriented PLMs. We release two new legal PLMs trained on LeXFiles and evaluate them alongside others on LegalLAMA and LexGLUE. We find that probing performance strongly correlates with upstream performance in related legal topics. On the other hand, downstream performance is mainly driven by the model's size and prior legal knowledge which can be estimated by upstream and probing performance. Based on these findings, we can conclude that both dimensions are important for those seeking the development of domain-specific PLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-TuningChunlin Tian, Zhan Shi, Zhijiang Guo, Li Li et al.NeurIPS 2024 · 172 citations
- Defining Knowledge: Bridging Epistemology and Large Language ModelsConstanza Fierro, Ruchira Dhar, Filippos Stamatiou, Nicolas Garneau et al.EMNLP 2024 · 5 citations
- ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract DescriptionsSreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Reddy Evuru et al.ACL 2024 · 3 citations
- Measuring the Effect of Disfluency in Multilingual Knowledge Probing BenchmarksKirill Semenov, Rico SennrichEMNLP 2025 · 2 citations
- ProMALex: Progressive Modular Adapters for Multi-Jurisdictional Legal Language ModelingT. Y. S. S. Santosh, Mohamed Hesham ElganayniACL 2025 · 2 citations
Builds on7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transferIlias Chalkidis, Manos Fergadiotis, Ion AndroutsopoulosEMNLP 2021 · 78 citations
- UNKs Everywhere: Adapting Multilingual Language Models to New ScriptsJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2021 · 3 citations
- ILDC for CJPE: Indian Legal Documents Corpus for Court Judgment Prediction and ExplanationVijit Malik, Rishabh Sanjay, Shubham Kumar Nigam, Kripabandhu Ghosh et al.ACL 2021
Related papers
- LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II et al.ACL 2022
- MultiLegalPile: A 689GB Multilingual Legal CorpusJoel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis et al.ACL 2024
- Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language ModelsZaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su et al.ACL 2022 · 36 citations
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou et al.EMNLP 2024 · 59 citations
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li et al.EMNLP 2022 · 27 citations
