MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer
Ilias Chalkidis, Manos Fergadiotis, Ion Androutsopoulos
摘要
We introduce MULTI-EURLEX, a new multilingual dataset for topic classification of legal documents. The dataset comprises 65k European Union (EU) laws, officially translated in 23 languages, annotated with multiple labels from the EUROVOC taxonomy. We highlight the effect of temporal concept drift and the importance of chronological, instead of random splits. We use the dataset as a testbed for zeroshot cross-lingual transfer, where we exploit annotated training documents in one language (source) to classify documents in another language (target). We find that fine-tuning a multilingually pretrained model (XLM-ROBERTA, MT5) in a single source language leads to catastrophic forgetting of multilingual knowledge and, consequently, poor zero-shot transfer to other languages. Adaptation strategies, namely partial fine-tuning, adapters, BITFIT, LNFIT, originally proposed to accelerate finetuning for new end-tasks, help retain multilingual knowledge from pretraining, substantially improving zero-shot cross-lingual transfer, but their impact also depends on the pretrained model used and the size of the label set. Language ISO Member Countries where official EU Speakers (%) Number of Documents Words per document code Native Total Train Dev. Test English en United Kingdom (1973-2020),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal DomainDennis Aumiller, Ashish Chouhan, Michael GertzEMNLP 2022 · 被引用 31 次
- LeXFiles and LegalLAMA: Facilitating English Multinational Legal Language Model DevelopmentIlias Chalkidis, Nicolas Garneau, Catalina Goanta, Daniel Martin Katz 等ACL 2023 · 被引用 29 次
- Knowledge Transfer in Incremental Learning for Multilingual Neural Machine TranslationKaiyu Huang, Peng Li, Jin Ma, Ting Yao 等ACL 2023 · 被引用 17 次
- MMTEB: Massive Multilingual Text Embedding BenchmarkKenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos 等ICLR 2025 · 被引用 10 次
- EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation LearningPing Guo, Xiangpeng Wei, Yue Hu, Baosong Yang 等NeurIPS 2023 · 被引用 8 次
它引用的顶会 Paper10
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
- CamemBERT: a Tasty French Language ModelLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont 等ACL 2020 · 被引用 703 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- How Does NLP Benefit Legal System: A Summary of Legal Artificial IntelligenceHaoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang 等ACL 2020 · 被引用 316 次
相关 Paper
- ZGUL: Zero-shot Generalization to Unseen Languages using Multi-source Ensembling of Language AdaptersVipul Rathore, Rajdeep Dhingra, Parag Singla, MausamEMNLP 2023
- Less-forgetting Multi-lingual Fine-tuningYuren Mao, Yaobo Liang, Nan Duan, Haobo Wang 等NeurIPS 2022 · 被引用 10 次
- Improving Zero-Shot Cross-Lingual Transfer Learning via Robust TrainingKuan-Hao Huang, Wasi Uddin Ahmad, Nanyun Peng, Kai-Wei ChangEMNLP 2021 · 被引用 29 次
- Model Selection for Cross-lingual TransferYang Chen, Alan RitterEMNLP 2021
- The Benefits of Label-Description Training for Zero-Shot Text ClassificationLingyu Gao, Debanjan Ghosh, Kevin GimpelEMNLP 2023 · 被引用 6 次
