LeXFiles and LegalLAMA: Facilitating English Multinational Legal Language Model Development
Ilias Chalkidis, Nicolas Garneau, Catalina Goanta, Daniel Martin Katz, Anders Søgaard
摘要
In this work, we conduct a detailed analysis on the performance of legal-oriented pre-trained language models (PLMs). We examine the interplay between their original objective, acquired knowledge, and legal language understanding capacities which we define as the upstream, probing, and downstream performance, respectively. We consider not only the models' size but also the pre-training corpora used as important dimensions in our study. To this end, we release a multinational English legal corpus (LeXFiles) and a legal knowledge probing benchmark (LegalLAMA) to facilitate training and detailed analysis of legal-oriented PLMs. We release two new legal PLMs trained on LeXFiles and evaluate them alongside others on LegalLAMA and LexGLUE. We find that probing performance strongly correlates with upstream performance in related legal topics. On the other hand, downstream performance is mainly driven by the model's size and prior legal knowledge which can be estimated by upstream and probing performance. Based on these findings, we can conclude that both dimensions are important for those seeking the development of domain-specific PLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-TuningChunlin Tian, Zhan Shi, Zhijiang Guo, Li Li 等NeurIPS 2024 · 被引用 172 次
- Defining Knowledge: Bridging Epistemology and Large Language ModelsConstanza Fierro, Ruchira Dhar, Filippos Stamatiou, Nicolas Garneau 等EMNLP 2024 · 被引用 5 次
- ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract DescriptionsSreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Reddy Evuru 等ACL 2024 · 被引用 3 次
- Measuring the Effect of Disfluency in Multilingual Knowledge Probing BenchmarksKirill Semenov, Rico SennrichEMNLP 2025 · 被引用 2 次
- ProMALex: Progressive Modular Adapters for Multi-Jurisdictional Legal Language ModelingT. Y. S. S. Santosh, Mohamed Hesham ElganayniACL 2025 · 被引用 2 次
它引用的顶会 Paper7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transferIlias Chalkidis, Manos Fergadiotis, Ion AndroutsopoulosEMNLP 2021 · 被引用 78 次
- UNKs Everywhere: Adapting Multilingual Language Models to New ScriptsJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2021 · 被引用 3 次
- ILDC for CJPE: Indian Legal Documents Corpus for Court Judgment Prediction and ExplanationVijit Malik, Rishabh Sanjay, Shubham Kumar Nigam, Kripabandhu Ghosh 等ACL 2021
相关 Paper
- LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II 等ACL 2022
- MultiLegalPile: A 689GB Multilingual Legal CorpusJoel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis 等ACL 2024
- Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language ModelsZaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su 等ACL 2022 · 被引用 36 次
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou 等EMNLP 2024 · 被引用 59 次
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li 等EMNLP 2022 · 被引用 27 次
