Detecting Legal Citations in United Kingdom Court Judgments
Holli Sargeant, Andreas Östling, Måns Magnusson
摘要
Legal citation detection in court judgments underpins reliable precedent mapping, citation analytics, and document retrieval. Extracting references to legislation and case law in the United Kingdom is especially challenging: citation styles have evolved over centuries, and judgments routinely cite foreign or historical authorities. We conduct the first systematic comparison of three modelling paradigms on this task using the Cambridge Law Corpus: (i) rule-based regular expressions; (ii) transformer-based encoders (BERT, RoBERTa, LEGAL-BERT, ModernBERT); and (iii) large language models (GPT-4.1). We produced a gold-standard high-quality corpus of 190 court judgments containing 45,179 fine-grained annotations for UK and non-UK legislation and case references. ModernBERT achieves a macroaveraged F1 of 93.3%, only marginally ahead of the other encoder-only models, yet significantly outperforming the strongest regularexpression baseline (35.42% F1) and GPT-4.1 (76.57% F1).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Lawma: The Power of Specialization for Legal AnnotationRicardo Dominguez-Olmedo, Vedant Nanda, Rediet Abebe, Stefan Bechtold 等ICLR 2025 · 被引用 1 次
- Modeling Legal Reasoning: LM Annotation at the Edge of Human AgreementRosamond Elizabeth Thalken, Edward H. Stiglitz, David Mimno, Matthew WilkensEMNLP 2023 · 被引用 9 次
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou 等EMNLP 2024 · 被引用 59 次
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller 等ACL 2025 · 被引用 552 次
- MultiLegalPile: A 689GB Multilingual Legal CorpusJoel Niklaus, Veton Matoshi, Matthias Stürmer, Ilias Chalkidis 等ACL 2024
