KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
Jihyung Lee, Daehui Kim, Seonjeong Hwang, Hyounghun Kim, Gary Lee
摘要
Large Language Models (LLM) have achieved remarkable performances in general domains and are now extending into the expert domain of law.Several benchmarks have been proposed to evaluate LLMs' legal capabilities.However, these benchmarks fail to evaluate open-ended and provisiongrounded Question Answering (QA).To address this, we introduce a Korean Benchmark for Legal EXplainable QA (KOBLEX), designed to evaluate provision-grounded, multihop legal reasoning.KOBLEX includes 226 scenario-based QA instances and their supporting provisions, created using a hybrid LLM-human expert pipeline.We also propose a method called Parametric provisionguided Selection Retrieval (PARSER), which uses LLM-generated parametric provisions to guide legally grounded and reliable answers.PARSER facilitates multi-hop reasoning on complex legal questions by generating parametric provisions and employing a three-stage sequential retrieval process.Furthermore, to better evaluate the legal fidelity of the generated answers, we propose Legal Fidelity Evaluation (LF-EVAL).LF-EVAL is an automatic metric that jointly considers the question, answer, and supporting provisions and shows a high correlation with human judgments.Experimental results show that PARSER consistently outperforms strong baselines, achieving the best results across multiple LLMs.Notably, compared to standard retrieval with GPT-4o, PARSER achieves 37.91 higher F-1 and 30.81 higher LF-EVAL.Further analyses reveal that PARSER efficiently delivers consistent performance across reasoning depths, with ablations confirming the effectiveness of PARSER. 1 * Equal Contribution. 1 The code and dataset are available at https://github. com/daehuikim/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
相关 Paper
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou 等EMNLP 2024 · 被引用 59 次
- PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal PracticeYuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song 等ACL 2026 · 被引用 7 次
- LEXam: Benchmarking Legal Reasoning on 340 Law ExamsYu Fan, Jingwei Ni, Jakob Merane, Yang Tian 等ICLR 2026 · 被引用 56 次
- Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative InferenceHua Cai, Shuang Zhao, Liang Zhang, Xuli Shen 等EMNLP 2025
- Evaluating Legal Reasoning Traces with Legal Issue Tree RubricsJinu Lee, Kyoung-Woon On, Sophia Simeng Han, Arman Cohan 等ACL 2026 · 被引用 2 次
