AGB-DE: A Corpus for the Automated Legal Assessment of Clauses in German Consumer Contracts
Daniel Braun, Florian Matthes
Abstract
Legal tasks and datasets are often used as benchmarks for the capabilities of language models. However, openly available annotated datasets are rare. In this paper, we introduce AGB-DE, a corpus of 3,764 clauses from German consumer contracts that have been annotated and legally assessed by legal experts. Together with the data, we present a first baseline for the task of detecting potentially void clauses, comparing the performance of an SVM baseline with three fine-tuned open language models and the performance of GPT-3.5. Our results show the challenging nature of the task, with no approach exceeding an F1-score of 0.54. While the fine-tuned models often performed better with regard to precision, GPT-3.5 outperformed the other approaches with regard to recall. An analysis of the errors indicates that one of the main challenges could be the correct interpretation of complex clauses, rather than the decision boundaries of what is permissible and what is not.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09d5abba-cf26-40c3-ad19-3798fb9c8dc1Cited by top-tier papers1
Ask how each one uses itBuilds on2
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II et al.ACL 2022
Related papers
- Agent-Specific Deontic Modality Detection in Legal LanguageAbhilasha Sancheti, Aparna Garimella, Balaji Vasan Srinivasan, Rachel RudingerEMNLP 2022 · 3 citations
- Lawma: The Power of Specialization for Legal AnnotationRicardo Dominguez-Olmedo, Vedant Nanda, Rediet Abebe, Stefan Bechtold et al.ICLR 2025 · 1 citation
- Validating Formal Specifications with LLM-Generated Test CasesAlcino Cunha, Nuno MacedoFM 2026 · 1 citation
- ACORD: An Expert-Annotated Retrieval Dataset for Legal Contract DraftingSteven H. Wang, Maksim Zubkov, Kexin Fan, Sarah Harrell et al.ACL 2025 · 14 citations
- Modeling Legal Reasoning: LM Annotation at the Edge of Human AgreementRosamond Elizabeth Thalken, Edward H. Stiglitz, David Mimno, Matthew WilkensEMNLP 2023 · 9 citations
