A Statutory Article Retrieval Dataset in French
Antoine Louis, Gerasimos Spanakis
Abstract
Statutory article retrieval is the task of automatically retrieving law articles relevant to a legal question. While recent advances in natural language processing have sparked considerable interest in many legal tasks, statutory article retrieval remains primarily untouched due to the scarcity of large-scale and high-quality annotated datasets. To address this bottleneck, we introduce the Belgian Statutory Article Retrieval Dataset (BSARD), which consists of 1,100+ French native legal questions labeled by experienced jurists with relevant articles from a corpus of 22,600+ Belgian law articles. Using BSARD, we benchmark several state-of-the-art retrieval approaches, including lexical and dense architectures, both in zero-shot and supervised setups. We find that fine-tuned dense retrieval models significantly outperform other systems. Our best performing baseline achieves 74.8% R@100, which is promising for the feasibility of the task and indicates there is still room for improvement. By the specificity of the domain and addressed task, BSARD presents a unique challenge problem for future research on legal information retrieval. Our dataset and source code are publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c61171c-ba98-460c-8a5e-59925cecb51fCited by top-tier papers9
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou et al.EMNLP 2024 · 59 citations
- Explicitly Integrating Judgment Prediction with Legal Document Retrieval: A Law-Guided Generative ApproachWeicong Qin, Zelin Cao, Weijie Yu, Zihua Si et al.SIGIR 2024 · 17 citations
- VendorLink: An NLP approach for Identifying & Linking Vendor Migrants & Potential Aliases on Darknet MarketsVageesh Saxena, Nils Rethmeier, Gijs van Dijck, Gerasimos SpanakisACL 2023 · 6 citations
- MAIR: A Massive Benchmark for Evaluating Instructed RetrievalWeiwei Sun, Zhengliang Shi, Wu Long, Lingyong Yan et al.EMNLP 2024 · 1 citation
- Triple-Encoders: Representations That Fire Together, Wire TogetherJustus-Jonas Erker, Florian Mai, Nils Reimers, Gerasimos Spanakis et al.ACL 2024
Builds on6
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- CamemBERT: a Tasty French Language ModelLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont et al.ACL 2020 · 703 citations
- Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep LearningHamza Harkous, Kassem Fawaz, Rémi Lebret, Florian Schaub et al.USENIX Security 2018 · 400 citations
- ETC: Encoding Long and Structured Inputs in TransformersJoshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek et al.EMNLP 2020 · 268 citations
- JEC-QA: A Legal-Domain Question Answering DatasetHaoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang et al.AAAI 2020 · 212 citations
Related papers
- LePaRD: A Large-Scale Dataset of Judicial Citations to PrecedentRobert Mahari, Dominik Stammbach, Elliott Ash, Alex PentlandACL 2024 · 1 citation
- Factoring Statutory Reasoning as Language Understanding ChallengesNils Holzenberger, Benjamin Van DurmeACL 2021
- IL-PCSR: Legal Corpus for Prior Case and Statute RetrievalShounak Paul, Dhananjay Ghumare, Pawan Goyal, Saptarshi Ghosh et al.EMNLP 2025
- LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements GenerationChaeeun Kim, Jinu Lee, Wonseok HwangEMNLP 2025 · 4 citations
- Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate PairsCheng Gao, Chaojun Xiao, Zhenghao Liu, Huimin Chen et al.EMNLP 2024 · 1 citation
