LexCLiPR: Cross-Lingual Paragraph Retrieval from Legal Judgments
Rohit Upadhya, T. Y. S. S. Santosh
Abstract
Efficient retrieval of pinpointed information from case law is crucial for legal professionals but challenging due to the length and complexity of legal judgments. Existing works mostly often focus on retrieving entire cases rather than precise, paragraph-level information. Moreover, multilingual legal practice necessitates cross-lingual retrieval, most works have been limited to monolingual settings. To address these gaps, we introduce LexCLiPR, a cross-lingual dataset for paragraph-level retrieval from European Court of Human Rights (ECtHR) judgments, leveraging multilingual case law guides and distant supervision to curate our dataset. We evaluate retrieval models in a zero-shot setting, revealing the limitations of pre-trained multilingual models for crosslingual tasks in low-resource languages and the importance of retrieval based post-training strategies. In fine-tuning settings, we observe that two-tower models excel in cross-lingual retrieval, while siamese architectures are better suited for monolingual tasks. Fine-tuning multilingual models on native language queries improves performance but struggles to generalize to unseen legal concepts, highlighting the need for robust strategies to address topical distribution shifts in the legal queries. 1 . * https://www.echr.coe.int/knowledge-sharing * Not all guides are available in every language *
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22d4c152-8564-4b7d-997d-4f4d1e1c0de5Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- LeSICiN: A Heterogeneous Graph-Based Approach for Automatic Legal Statute Identification from Indian Legal DocumentsShounak Paul, Pawan Goyal, Saptarshi GhoshAAAI 2022 · 61 citations
- CLIRMatrix: A massively large collection of bilingual and multilingual datasets for Cross-Lingual Information RetrievalShuo Sun, Kevin DuhEMNLP 2020 · 53 citations
Related papers
- EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal DomainDennis Aumiller, Ashish Chouhan, Michael GertzEMNLP 2022 · 31 citations
- Cross-lingual Language Model Pretraining for RetrievalPuxuan Yu, Hongliang Fei, Ping LiWWW 2021 · 42 citations
- MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transferIlias Chalkidis, Manos Fergadiotis, Ion AndroutsopoulosEMNLP 2021 · 78 citations
- EUROPA: A Legal Multilingual Keyphrase Generation DatasetOlivier Salaün, Frédéric Piedboeuf, Guillaume Le Berre, David Alfonso-Hermelo et al.ACL 2024
- MultiLegalPile: A 689GB Multilingual Legal CorpusJoel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis et al.ACL 2024
