PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
Yuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song, Xinran Xu, Yubo Ma, Tianyi Tang, Li Zhang, Qingjing Chen, Di Feng, Wenbo Lv, Weiheng Wu
Abstract
As large language models (LLMs) are increasingly applied to legal domain-specific tasks, evaluating their ability to perform legal work in real-world settings has become essential. However, existing legal benchmarks rely on simplified and highly standardized tasks, failing to capture the ambiguity, complexity, and reasoning demands of real legal practice. Moreover, prior evaluations often adopt coarse, singledimensional metrics and do not explicitly assess fine-grained legal reasoning. To address these limitations, we introduce PLAWBENCH, a Practical Law Benchmark designed to evaluate LLMs in realistic legal practice scenarios. Grounded in real-world legal workflows, PLAWBENCH models the core processes of legal practitioners through three task categories: public legal consultation, practical case analysis, and legal document generation. These tasks assess a model's ability to identify legal issues and key facts, perform structured legal reasoning, and generate legally coherent documents. PLAWBENCH comprises 850 questions across 13 practical legal scenarios, with each question accompanied by expert-designed evaluation rubrics, resulting in approximately 12,500 rubric items for fine-grained assessment. Using an LLM-based evaluator aligned with human expert judgments, we evaluate 10 state-ofthe-art LLMs. Experimental results show that none achieves strong performance on PLAW-BENCH, revealing substantial limitations in the fine-grained legal reasoning capabilities of current LLMs and highlighting important directions for future evaluation and development of legal LLMs. Data is available at: https: //github.com/skylenage/PLawbench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9bce962-4eb8-4b52-bdd6-52ca67a34f5cBuilds on2
Related papers
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou et al.EMNLP 2024 · 59 citations
- Pub-LawBench: Public-Oriented Benchmarking for LegalAIQiaoyu Zheng, Zehan Ma, Yijing Zhang, Qiqi Wang et al.ACL 2026
- PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional ReasoningAfra Feyza Akyürek, Advait Gosai, Chen Bo Calvin Zhang, Vipul Gupta et al.ACL 2026 · 18 citations
- LegalAgentBench: Evaluating LLM Agents in Legal DomainHaitao Li, Junjie Chen, Jingli Yang, Qingyao Ai et al.ACL 2025
- Measuring the Unmeasurable: Unveiling Latent Cognitive Capabilities of LLMCui Danxin, Sihang Jiang, Keyi Wang, Zhiyi Duan et al.AAAI 2026
