Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate Pairs
Cheng Gao, Chaojun Xiao, Zhenghao Liu, Huimin Chen, Zhiyuan Liu, Maosong Sun
Abstract
Legal case retrieval (LCR) aims to provide similar cases as references for a given fact description. This task is crucial for promoting consistent judgments in similar cases, effectively enhancing judicial fairness and improving work efficiency for judges. However, existing works face two main challenges for real-world applications: existing works mainly focus on case-to-case retrieval using lengthy queries, which does not match real-world scenarios; and the limited data scale, with current datasets containing only hundreds of queries, is insufficient to satisfy the training requirements of existing data-hungry neural models. To address these issues, we introduce an automated method to construct synthetic query-candidate pairs and build the largest LCR dataset to date, LEAD, which is hundreds of times larger than existing datasets. This data construction method can provide ample training signals for LCR models. Experimental results demonstrate that model training with our constructed data can achieve state-of-the-art results on two widely-used LCR benchmarks. Besides, the construction method can also be applied to civil cases and achieve promising results. The data and codes can be found in https://github.com/thunlp/LEAD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 066524a9-60ec-48a2-82a7-211d64562fa8Cited by top-tier papers1
Ask how each one uses itBuilds on11
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Pre-training Tasks for Embedding-based Large-scale RetrievalWei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang et al.ICLR 2020 · 325 citations
- End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question AnsweringDevendra Singh Sachan, Siva Reddy, William L. Hamilton, Chris Dyer et al.NeurIPS 2021 · 197 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Legal Case Retrieval: A Survey of the State of the ArtYi Feng, Chuanyi Li, Vincent NgACL 2024 · 15 citations
- LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements GenerationChaeeun Kim, Jinu Lee, Wonseok HwangEMNLP 2025 · 4 citations
- Learning Interpretable Legal Case Retrieval via Knowledge-Guided Case ReformulationChenlong Deng, Kelong Mao, Zhicheng DouEMNLP 2024 · 3 citations
- Unsupervised Legal Evidence Retrieval via Contrastive Learning with Approximate Aggregated PositiveFeng Yao, Jingyuan Zhang, Yating Zhang, Xiaozhong Liu et al.AAAI 2023 · 9 citations
- CFGL-LCR: A Counterfactual Graph Learning Framework for Legal Case RetrievalKun Zhang, Chong Chen, Yuanzhuo Wang, Qi Tian et al.KDD 2023 · 9 citations
