From Dissonance to Insights: Dissecting Disagreements in Rationale Construction for Case Outcome Classification
Shanshan Xu, T. Y. S. S. Santosh, Oana Ichim, Isabella Risini, Barbara Plank, Matthias Grabmair
Abstract
In legal NLP, Case Outcome Classification<br/>(COC) must not only be accurate but also<br/>trustworthy and explainable. Existing work<br/>in explainable COC has been limited to an-<br/>notations by a single expert. However, it is<br/>well-known that lawyers may disagree in their<br/>assessment of case facts. We hence collect<br/>a novel dataset RAVE: Rationale Variation<br/>in ECHR1, which is obtained from two ex-<br/>perts in the domain of international human<br/>rights law, for whom we observe weak agree-<br/>ment. We study their disagreements and build a<br/>two-level task-independent taxonomy, supple-<br/>mented with COC-specific subcategories. We<br/>quantitatively assess different taxonomy cate-<br/>gories and find that disagreements mainly stem<br/>from underspecification of the legal context,<br/>which poses challenges given the typically lim-<br/>ited granularity and noise in COC metadata. To<br/>our knowledge, this is the first work in the legal<br/>NLP that focuses on building a taxonomy over<br/>human label variation. We further assess the ex-<br/>plainablility of state-of-the-art COC models on<br/>RAVE and observe limited agreement between<br/>models and experts. Overall, our case study re-<br/>veals hitherto underappreciated complexities in<br/>creating benchmark datasets in legal NLP that<br/>revolve around identifying aspects of a case’s<br/>facts supposedly relevant to its outcome
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ea26cbc-2feb-452a-902c-c3cd70ffcd83Cited by top-tier papers2
- Which Demographics do LLMs Default to During Annotation?Johannes Schäfer, Aidan Combs, Christopher Bagdon, Jiahui Li et al.ACL 2025 · 11 citations
- Through the Lens of Split Vote: Exploring Disagreement, Difficulty and Calibration in Legal Case Outcome ClassificationShanshan Xu, T. Y. S. S. Santosh, Oana Ichim, Barbara Plank et al.ACL 2024
Builds on4
- NeurJudge: A Circumstance-aware Neural Framework for Legal Judgment PredictionLinan Yue, Qi Liu, Binbin Jin, Han Wu et al.SIGIR 2021 · 83 citations
- Deconfounding Legal Judgment Prediction for European Court of Human Rights Cases Towards Better Alignment with ExpertsTokala Yaswanth Sri Sai Santosh, Shanshan Xu, Oana Ichim, Matthias GrabmairEMNLP 2022 · 13 citations
- ILDC for CJPE: Indian Legal Documents Corpus for Court Judgment Prediction and ExplanationVijit Malik, Rishabh Sanjay, Shubham Kumar Nigam, Kripabandhu Ghosh et al.ACL 2021
- LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II et al.ACL 2022
Related papers
- Explainable Legal Case Matching via Inverse Optimal Transport-based Rationale ExtractionWeijie Yu, Zhongxiang Sun, Jun Xu, Zhenhua Dong et al.SIGIR 2022 · 45 citations
- Evaluating Legal Reasoning Traces with Legal Issue Tree RubricsJinu Lee, Kyoung-Woon On, Sophia Simeng Han, Arman Cohan et al.ACL 2026 · 2 citations
- CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative StructuresEkaterina Sviridova, Anar Yeginbergen, Ainara Estarrona, Elena Cabrio et al.EMNLP 2024 · 2 citations
- e-CARE: a New Dataset for Exploring Explainable Causal ReasoningLi Du, Xiao Ding, Kai Xiong, Ting Liu et al.ACL 2022
- Legal Fact Prediction: The Missing Piece in Legal Judgment PredictionJunkai Liu, Yujie Tong, Hui Huang, Bowen Zheng et al.EMNLP 2025
