DELTA: Pre-Train a Discriminative Encoder for Legal Case Retrieval via Structural Word Alignment
Haitao Li, Qingyao Ai, Xinyan Han, Jia Chen, Qian Dong, Yiqun Liu
Abstract
Recent research demonstrates the effectiveness of using pre-trained language models for legal case retrieval. Most of the existing works focus on improving the representation ability for the contextualized embedding of the [CLS] token and calculate relevance using textual semantic similarity. However, in the legal domain, textual semantic similarity does not always imply that the cases are relevant enough. Instead, relevance in legal cases primarily depends on the similarity of key facts that impact the final judgment. Without proper treatments, the discriminative ability of learned representations could be limited since legal cases are lengthy and contain numerous non-key facts. To this end, we introduce DELTA, a discriminative model designed for legal case retrieval. The basic idea involves pinpointing key facts in legal cases and pulling the contextualized embedding of the [CLS] token closer to the key facts while pushing away from the non-key facts, which can warm up the case embedding space in an unsupervised manner. To be specific, this study brings the word alignment mechanism to the contextual masked auto-encoder. First, we leverage shallow decoders to create information bottlenecks, aiming to enhance the representation ability. Second, we employ the deep decoder to enable ``translation'' between different structures, with the goal of pinpointing key facts to enhance discriminative ability. Comprehensive experiments conducted on publicly available legal benchmarks show that our approach can outperform existing state-of-the-art methods in legal case retrieval. It provides a new perspective on the in-depth understanding and processing of legal case documents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89ce3ed5-7724-46d3-83d4-8dcda69cf5c5Builds on6
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- SAILER: Structure-aware Pre-trained Language Model for Legal Case RetrievalHaitao Li, Qingyao Ai, Jia Chen, Qian Dong et al.SIGIR 2023 · 68 citations
- ConTextual Masked Auto-Encoder for Dense Passage RetrievalXing Wu, Guangyuan Ma, Meng Lin, Zijia Lin et al.AAAI 2023 · 34 citations
- CaseEncoder: A Knowledge-enhanced Pre-trained Model for Legal Case EncodingYixiao Ma, Yueyue Wu, Weihang Su, Qingyao Ai et al.EMNLP 2023 · 5 citations
Related papers
- CFGL-LCR: A Counterfactual Graph Learning Framework for Legal Case RetrievalKun Zhang, Chong Chen, Yuanzhuo Wang, Qi Tian et al.KDD 2023 · 9 citations
- Unsupervised Legal Evidence Retrieval via Contrastive Learning with Approximate Aggregated PositiveFeng Yao, Jingyuan Zhang, Yating Zhang, Xiaozhong Liu et al.AAAI 2023 · 9 citations
- Learning Interpretable Legal Case Retrieval via Knowledge-Guided Case ReformulationChenlong Deng, Kelong Mao, Zhicheng DouEMNLP 2024 · 3 citations
- CaseLink: Inductive Graph Learning for Legal Case RetrievalYanran Tang, Ruihong Qiu, Hongzhi Yin, Xue Li et al.SIGIR 2024 · 10 citations
- Law Article-Enhanced Legal Case Matching: A Causal Learning ApproachZhongxiang Sun, Jun Xu, Xiao Zhang, Zhenhua Dong et al.SIGIR 2023 · 25 citations
