DELTA: Pre-Train a Discriminative Encoder for Legal Case Retrieval via Structural Word Alignment
Haitao Li, Qingyao Ai, Xinyan Han, Jia Chen, Qian Dong, Yiqun Liu
摘要
Recent research demonstrates the effectiveness of using pre-trained language models for legal case retrieval. Most of the existing works focus on improving the representation ability for the contextualized embedding of the [CLS] token and calculate relevance using textual semantic similarity. However, in the legal domain, textual semantic similarity does not always imply that the cases are relevant enough. Instead, relevance in legal cases primarily depends on the similarity of key facts that impact the final judgment. Without proper treatments, the discriminative ability of learned representations could be limited since legal cases are lengthy and contain numerous non-key facts. To this end, we introduce DELTA, a discriminative model designed for legal case retrieval. The basic idea involves pinpointing key facts in legal cases and pulling the contextualized embedding of the [CLS] token closer to the key facts while pushing away from the non-key facts, which can warm up the case embedding space in an unsupervised manner. To be specific, this study brings the word alignment mechanism to the contextual masked auto-encoder. First, we leverage shallow decoders to create information bottlenecks, aiming to enhance the representation ability. Second, we employ the deep decoder to enable ``translation'' between different structures, with the goal of pinpointing key facts to enhance discriminative ability. Comprehensive experiments conducted on publicly available legal benchmarks show that our approach can outperform existing state-of-the-art methods in legal case retrieval. It provides a new perspective on the in-depth understanding and processing of legal case documents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- SAILER: Structure-aware Pre-trained Language Model for Legal Case RetrievalHaitao Li, Qingyao Ai, Jia Chen, Qian Dong 等SIGIR 2023 · 被引用 68 次
- ConTextual Masked Auto-Encoder for Dense Passage RetrievalXing Wu, Guangyuan Ma, Meng Lin, Zijia Lin 等AAAI 2023 · 被引用 34 次
- CaseEncoder: A Knowledge-enhanced Pre-trained Model for Legal Case EncodingYixiao Ma, Yueyue Wu, Weihang Su, Qingyao Ai 等EMNLP 2023 · 被引用 5 次
相关 Paper
- CFGL-LCR: A Counterfactual Graph Learning Framework for Legal Case RetrievalKun Zhang, Chong Chen, Yuanzhuo Wang, Qi Tian 等KDD 2023 · 被引用 9 次
- Unsupervised Legal Evidence Retrieval via Contrastive Learning with Approximate Aggregated PositiveFeng Yao, Jingyuan Zhang, Yating Zhang, Xiaozhong Liu 等AAAI 2023 · 被引用 9 次
- Learning Interpretable Legal Case Retrieval via Knowledge-Guided Case ReformulationChenlong Deng, Kelong Mao, Zhicheng DouEMNLP 2024 · 被引用 3 次
- CaseLink: Inductive Graph Learning for Legal Case RetrievalYanran Tang, Ruihong Qiu, Hongzhi Yin, Xue Li 等SIGIR 2024 · 被引用 10 次
- Law Article-Enhanced Legal Case Matching: A Causal Learning ApproachZhongxiang Sun, Jun Xu, Xiao Zhang, Zhenhua Dong 等SIGIR 2023 · 被引用 25 次
