Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
Zheng Jia, Shengbin Yue, Wei Chen, Siyuan Wang, Yidong Liu, Zejun Li, Yun Song, Zhongyu Wei
摘要
The gap between existing benchmarks and the dynamic nature of real-world legal practice poses a key barrier to advancing legal intelligence. To this end, we introduce J1-ENVS, the first interactive and dynamic legal environment tailored for LLM-based agents. Guided by legal experts, it comprises six representative scenarios from Chinese legal practices at three levels of environmental complexity. We further introduce J1-EVAL, a dual-metric evaluation framework, designed to assess both task performance and procedural compliance across varying levels of legal proficiency. Extensive experiments on 17 LLM agents reveal that while many models demonstrate solid legal knowledge, they struggle with procedural execution in dynamic settings. Even the SOTA model is below 60% overall performance Ṫhese findings highlight persistent challenges in achieving dynamic legal intelligence and offer valuable insights to guide future research. Resources will be available upon acceptance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LLM Agents in Law: Taxonomy, Applications, and ChallengesShuang Liu, Ruijia Zhang, Ruoyun Ma, Yujia Deng 等ACL 2026 · 被引用 3 次
- CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive EngagementHong Qian, Yuanhao Liu, Zihan Zhou, Zongbao Zhang 等ICML 2026
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- JEC-QA: A Legal-Domain Question Answering DatasetHaoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang 等AAAI 2020 · 被引用 212 次
- Syllogistic Reasoning for Legal Judgment AnalysisWentao Deng, Jiahuan Pei, Keyi Kong, Zhe Chen 等EMNLP 2023 · 被引用 17 次
- LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II 等ACL 2022
相关 Paper
- LegalAgentBench: Evaluating LLM Agents in Legal DomainHaitao Li, Junjie Chen, Jingli Yang, Qingyao Ai 等ACL 2025
- JurisBench: A Deep Benchmark for Assessing Large Language Models in Professional Legal PracticeZiang Chen, Guannan Li, Fanlin Ji, Yipeng Kang 等ACL 2026
- From Query to Counsel: Structured Reasoning with a Multi-Agent Framework and Dataset for Legal ConsultationMingfei Lu, Yi Zhang, Mengjia Wu, Yue FengACL 2026 · 被引用 2 次
- FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and ReasoningLiang Hu, Jianpeng Jiao, Jiashuo Liu, Dongyuan Mutu 等ICLR 2026 · 被引用 29 次
- PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal PracticeYuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song 等ACL 2026 · 被引用 7 次
