Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents
Ratnadira Widyasari, Martin Weyssow, Ivana Clairine Irsan, Han Wei Ang, Frank Liauw, Eng Lieh Ouh, Lwin Khin Shar, Hong Jin Kang, David Lo
Abstract
Detecting vulnerabilities in source code remains a critical yet challenging task, especially when benign and vulnerable functions share significant similarities. In this work, we introduce VulTrial, a courtroom-inspired multi-agent framework designed to identify vulnerable code and to provide explanations. It employs four role-specific agents, which are security researcher, code author, moderator, and review board. Using GPT-4o as the base LLM, VulTrial almost doubles the efficacy of prior best-performing baselines. Additionally, we show that role-specific instruction tuning with small quantities of data significantly further boosts VulTrial's efficacy. Our extensive experiments demonstrate the efficacy of VulTrial across different LLMs, including an open-source, in-house-deployable model (LLaMA-3.1-8B), as well as the high quality of its generated explanations and its ability to uncover multiple confirmed zero-day vulnerabilities in the wild.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b1b17fa-90d7-4bcc-8082-93726aa5d051Cited by top-tier papers4
- Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM AgentsPengfei He, Ash Fox, Lesly Miculicich, Stefan Friedli et al.ICML 2026 · 10 citations
- SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing ScenariosJunkai Chen, Huihui Huang, Yunbo Lyu, Junwen An et al.ACL 2026 · 5 citations
- Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability DetectionXin Peng, Bo Lin, Jing Wang, Xiaoling Li et al.FSE 2026 · 1 citation
- VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security SpecificationsHao Zhu, Jia Li, Cuiyun Gao, Jiaru Qian et al.FSE 2026
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
Related papers
- From Large to Mammoth: A Comparative Evaluation of Large Language Models in Vulnerability DetectionJie Lin, David MohaisenNDSS 2025
- Combining Fine-Tuning and LLM-Based Agents for Intuitive Smart Contract Auditing with JustificationsWei Ma, Daoyuan Wu, Yuqiang Sun, Tianwen Wang et al.ICSE 2025 · 28 citations
- Enhancing Vulnerability Detection via Inter-procedural Semantic CompletionBozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao et al.ISSTA 2025 · 2 citations
- Smart-LLaMA-DPO: Reinforced Large Language Model for Explainable Smart Contract Vulnerability DetectionLei Yu, Zhirong Huang, Hang Yuan, Shiqi Cheng et al.ISSTA 2025 · 13 citations
- CVE-Genie: An LLM-Based Multi-Agent Framework for Reproducing CVEsSaad Ullah, Praneeth Balasubramanian, Wenbo Guo, Amanda Burnett et al.CCS 2026
