KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
Jihyung Lee, Daehui Kim, Seonjeong Hwang, Hyounghun Kim, Gary Lee
Abstract
Large Language Models (LLM) have achieved remarkable performances in general domains and are now extending into the expert domain of law.Several benchmarks have been proposed to evaluate LLMs' legal capabilities.However, these benchmarks fail to evaluate open-ended and provisiongrounded Question Answering (QA).To address this, we introduce a Korean Benchmark for Legal EXplainable QA (KOBLEX), designed to evaluate provision-grounded, multihop legal reasoning.KOBLEX includes 226 scenario-based QA instances and their supporting provisions, created using a hybrid LLM-human expert pipeline.We also propose a method called Parametric provisionguided Selection Retrieval (PARSER), which uses LLM-generated parametric provisions to guide legally grounded and reliable answers.PARSER facilitates multi-hop reasoning on complex legal questions by generating parametric provisions and employing a three-stage sequential retrieval process.Furthermore, to better evaluate the legal fidelity of the generated answers, we propose Legal Fidelity Evaluation (LF-EVAL).LF-EVAL is an automatic metric that jointly considers the question, answer, and supporting provisions and shows a high correlation with human judgments.Experimental results show that PARSER consistently outperforms strong baselines, achieving the best results across multiple LLMs.Notably, compared to standard retrieval with GPT-4o, PARSER achieves 37.91 higher F-1 and 30.81 higher LF-EVAL.Further analyses reveal that PARSER efficiently delivers consistent performance across reasoning depths, with ablations confirming the effectiveness of PARSER. 1 * Equal Contribution. 1 The code and dataset are available at https://github. com/daehuikim/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
Related papers
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou et al.EMNLP 2024 · 59 citations
- PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal PracticeYuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song et al.ACL 2026 · 7 citations
- LEXam: Benchmarking Legal Reasoning on 340 Law ExamsYu Fan, Jingwei Ni, Jakob Merane, Yang Tian et al.ICLR 2026 · 56 citations
- Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative InferenceHua Cai, Shuang Zhao, Liang Zhang, Xuli Shen et al.EMNLP 2025
- Evaluating Legal Reasoning Traces with Legal Issue Tree RubricsJinu Lee, Kyoung-Woon On, Sophia Simeng Han, Arman Cohan et al.ACL 2026 · 2 citations
