MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer
Ilias Chalkidis, Manos Fergadiotis, Ion Androutsopoulos
Abstract
We introduce MULTI-EURLEX, a new multilingual dataset for topic classification of legal documents. The dataset comprises 65k European Union (EU) laws, officially translated in 23 languages, annotated with multiple labels from the EUROVOC taxonomy. We highlight the effect of temporal concept drift and the importance of chronological, instead of random splits. We use the dataset as a testbed for zeroshot cross-lingual transfer, where we exploit annotated training documents in one language (source) to classify documents in another language (target). We find that fine-tuning a multilingually pretrained model (XLM-ROBERTA, MT5) in a single source language leads to catastrophic forgetting of multilingual knowledge and, consequently, poor zero-shot transfer to other languages. Adaptation strategies, namely partial fine-tuning, adapters, BITFIT, LNFIT, originally proposed to accelerate finetuning for new end-tasks, help retain multilingual knowledge from pretraining, substantially improving zero-shot cross-lingual transfer, but their impact also depends on the pretrained model used and the size of the label set. Language ISO Member Countries where official EU Speakers (%) Number of Documents Words per document code Native Total Train Dev. Test English en United Kingdom (1973-2020),
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57accc48-c66d-4810-86d5-972313eae1a8Cited by top-tier papers13
- EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal DomainDennis Aumiller, Ashish Chouhan, Michael GertzEMNLP 2022 · 31 citations
- LeXFiles and LegalLAMA: Facilitating English Multinational Legal Language Model DevelopmentIlias Chalkidis, Nicolas Garneau, Catalina Goanta, Daniel Martin Katz et al.ACL 2023 · 29 citations
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao et al.ACL 2023 · 17 citations
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos et al.ICLR 2025 · 10 citations
- EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation LearningPing Guo, Xiangpeng Wei, Yue Hu, Baosong Yang et al.NeurIPS 2023 · 8 citations
Builds on10
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig et al.ICML 2020 · 1,132 citations
- CamemBERT: a Tasty French Language ModelLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont et al.ACL 2020 · 703 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- How Does NLP Benefit Legal System: A Summary of Legal Artificial IntelligenceHaoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang et al.ACL 2020 · 316 citations
Related papers
- ZGUL: Zero-shot Generalization to Unseen Languages using Multi-source Ensembling of Language AdaptersVipul Rathore, Rajdeep Dhingra, Parag Singla, MausamEMNLP 2023
- Less-forgetting Multi-lingual Fine-tuningYuren Mao, Yaobo Liang, Nan Duan, Haobo Wang et al.NeurIPS 2022 · 10 citations
- Improving Zero-Shot Cross-Lingual Transfer Learning via Robust TrainingKuan-Hao Huang, Wasi Uddin Ahmad, Nanyun Peng, Kai-Wei ChangEMNLP 2021 · 29 citations
- Model Selection for Cross-lingual TransferYang Chen, Alan RitterEMNLP 2021
- The Benefits of Label-Description Training for Zero-Shot Text ClassificationLingyu Gao, Debanjan Ghosh, Kevin GimpelEMNLP 2023 · 6 citations
