Pub-LawBench: Public-Oriented Benchmarking for LegalAI
Qiaoyu Zheng, Zehan Ma, Yijing Zhang, Qiqi Wang, Huijia Li, Qian Liu
摘要
Large language models (LLMs) are playing an increasingly pivotal role in LegalAI. However, existing benchmarks are primarily tailored for legal professionals, emphasizing deep reasoning and explainability. While public-facing legal applications demand outputs that are direct, actionable, and accessible, a need largely overlooked by current evaluation frameworks. To bridge this gap, we propose a public-oriented LegalAI benchmark grounded in legal functionalism and genre analysis. Specifically, we categorize public legal demands into two core tasks: Instant Question Answering and Legal Text Generation. We further introduce three public-oriented evaluation dimensions: legal normativity, content relevance, and format usability, which collectively assess the practical validity and user readiness of model outputs. To reflect real-world lay user usage, we evaluate 17 LLMs on Pub-LawBench using only simple prompts and Chain-of-Thought under a vanilla inference setting, excluding complex techniques like RAG or agent-based methods inaccessible to non-experts. Experiments reveal limitations of current LLMs in delivering effective public-oriented legal assistance, highlighting the need for more user-centric model development and benchmarking. 1 * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- JEC-QA: A Legal-Domain Question Answering DatasetHaoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang 等AAAI 2020 · 被引用 212 次
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou 等EMNLP 2024 · 被引用 59 次
- Precedent-Enhanced Legal Judgment Prediction with LLM and Domain-Model CollaborationYiquan Wu, Siying Zhou, Yifei Liu, Weiming Lu 等EMNLP 2023 · 被引用 37 次
相关 Paper
- PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal PracticeYuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song 等ACL 2026 · 被引用 7 次
- Expectation Alignment of Language Models for Real-World User ExpectationsMiaomiao Li, Yang Wang, Bin Liang, Shudong Liu 等ICML 2026
- LegalAgentBench: Evaluating LLM Agents in Legal DomainHaitao Li, Junjie Chen, Jingli Yang, Qingyao Ai 等ACL 2025
- ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language ModelsYuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang 等ACL 2024
- A User-Centric Multi-Intent Benchmark for Evaluating Large Language ModelsJiayin Wang, Fengran Mo, Weizhi Ma, Peijie Sun 等EMNLP 2024 · 被引用 10 次
