PAKTON: A Multi-Agent Framework for Question Answering in Long Legal Agreements
Petros Raptopoulos, Giorgos Filandrianos, Maria Lymperaiou, Giorgos Stamou
Abstract
Contract review is a complex and timeintensive task that typically demands specialized legal expertise, rendering it largely inaccessible to non-experts. Moreover, legal interpretation is rarely straightforward-ambiguity is pervasive, and judgments often hinge on subjective assessments. Compounding these challenges, contracts are usually confidential, restricting their use with proprietary models and necessitating reliance on open-source alternatives. To address these challenges, we introduce PAKTON: a fully open-source, endto-end, multi-agent framework with plug-andplay capabilities. PAKTON is designed to handle the complexities of contract analysis through collaborative agent workflows and a novel multi-stage retrieval-augmented generation (RAG) component, enabling automated legal document review that is more accessible, adaptable, and privacy-preserving. Experiments demonstrate that PAKTON outperforms both general-purpose and pretrained models in predictive accuracy, retrieval performance, explainability, completeness, and grounded justifications as evaluated through a human study and validated with automated metrics. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e84613e-45b0-4ef5-b96d-b03aa382dbefCited by top-tier papers1
Ask how each one uses itBuilds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
Related papers
- ProvBench: A Benchmark of Legal Provision Recommendation for Contract Auto-ReviewingXiuxuan Shen, Zhongyuan Jiang, Junsan Zhang, Junxiao Han et al.ACL 2025 · 2 citations
- ACORD: An Expert-Annotated Retrieval Dataset for Legal Contract DraftingSteven H. Wang, Maksim Zubkov, Kexin Fan, Sarah Harrell et al.ACL 2025 · 14 citations
- LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal ReasoningZerui Chen, Qinggang Zhang, Zhishang Xiang, Zhimin Wei et al.ACL 2026 · 2 citations
- DocETL: Agentic Query Rewriting and Evaluation for Complex Document ProcessingShreya Shankar, Tristan Chambers, Tarak Shah, Aditya G. Parameswaran et al.VLDB 2025 · 62 citations
- Automating Legal Interpretation with LLMs: Retrieval, Generation, and EvaluationKangcheng Luo, Quzhe Huang, Cong Jiang, Yansong FengACL 2025 · 4 citations
