A Statutory Article Retrieval Dataset in French
Antoine Louis, Gerasimos Spanakis
摘要
Statutory article retrieval is the task of automatically retrieving law articles relevant to a legal question. While recent advances in natural language processing have sparked considerable interest in many legal tasks, statutory article retrieval remains primarily untouched due to the scarcity of large-scale and high-quality annotated datasets. To address this bottleneck, we introduce the Belgian Statutory Article Retrieval Dataset (BSARD), which consists of 1,100+ French native legal questions labeled by experienced jurists with relevant articles from a corpus of 22,600+ Belgian law articles. Using BSARD, we benchmark several state-of-the-art retrieval approaches, including lexical and dense architectures, both in zero-shot and supervised setups. We find that fine-tuned dense retrieval models significantly outperform other systems. Our best performing baseline achieves 74.8% R@100, which is promising for the feasibility of the task and indicates there is still room for improvement. By the specificity of the domain and addressed task, BSARD presents a unique challenge problem for future research on legal information retrieval. Our dataset and source code are publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou 等EMNLP 2024 · 被引用 59 次
- Explicitly Integrating Judgment Prediction with Legal Document Retrieval: A Law-Guided Generative ApproachWeicong Qin, Zelin Cao, Weijie Yu, Zihua Si 等SIGIR 2024 · 被引用 17 次
- VendorLink: An NLP approach for Identifying & Linking Vendor Migrants & Potential Aliases on Darknet MarketsVageesh Saxena, Nils Rethmeier, Gijs van Dijck, Gerasimos SpanakisACL 2023 · 被引用 6 次
- MAIR: A Massive Benchmark for Evaluating Instructed RetrievalWeiwei Sun, Zhengliang Shi, Wu Long, Lingyong Yan 等EMNLP 2024 · 被引用 1 次
- Triple-Encoders: Representations That Fire Together, Wire TogetherJustus-Jonas Erker, Florian Mai, Nils Reimers, Gerasimos Spanakis 等ACL 2024
它引用的顶会 Paper6
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- CamemBERT: a Tasty French Language ModelLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont 等ACL 2020 · 被引用 703 次
- Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep LearningHamza Harkous, Kassem Fawaz, Rémi Lebret, Florian Schaub 等USENIX Security 2018 · 被引用 400 次
- ETC: Encoding Long and Structured Inputs in TransformersJoshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek 等EMNLP 2020 · 被引用 268 次
- JEC-QA: A Legal-Domain Question Answering DatasetHaoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang 等AAAI 2020 · 被引用 212 次
相关 Paper
- LePaRD: A Large-Scale Dataset of Judicial Citations to PrecedentRobert Mahari, Dominik Stammbach, Elliott Ash, Alex PentlandACL 2024 · 被引用 1 次
- Factoring Statutory Reasoning as Language Understanding ChallengesNils Holzenberger, Benjamin Van DurmeACL 2021
- IL-PCSR: Legal Corpus for Prior Case and Statute RetrievalShounak Paul, Dhananjay Ghumare, Pawan Goyal, Saptarshi Ghosh 等EMNLP 2025
- LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements GenerationChaeeun Kim, Jinu Lee, Wonseok HwangEMNLP 2025 · 被引用 4 次
- Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate PairsCheng Gao, Chaojun Xiao, Zhenghao Liu, Huimin Chen 等EMNLP 2024 · 被引用 1 次
