Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive Filtering
Yunpeng Xiong, Ting Zhang
Abstract
Static Application Security Testing (SAST) tools are essential for identifying software vulnerabilities, but they often produce a high volume of False Positives (FPs), imposing a substantial manual triage burden on developers. Recent advances in Large Language Model (LLM) agents offer a promising direction by enabling iterative reasoning, tool use, and environment interaction to refine SAST alerts. However, the comparative effectiveness of different LLM-based agent architectures for FP filtering remains poorly understood.
In this paper, we present a comparative study of three state-of-the-art LLM-based agent frameworks, i.e., Aider, OpenHands, and SWE-agent, for vulnerability FP filtering. We evaluate these frameworks using the vulnerabilities from the OWASP Benchmark and real-world open-source Java projects. We further conduct a focused post-cutoff C/C++ study using the strongest configuration to test contamination-free generalization and isolate key agentic capabilities. The experimental results show that LLM-based agents can remove the majority of SAST noise, reducing an initial FP detection rate of over 92% on the OWASP Benchmark to as low as 6.3% in the best configuration. On a real-world Java dataset, the best configuration of LLM-based agents can achieve an FP identification rate of up to 93.3% involving CodeQL alerts. However, the benefits of agents are strongly backbone-and CWE-dependent: agentic frameworks significantly outperform vanilla prompting for stronger models such as Claude Sonnet 4 and GPT-5, but yield limited or inconsistent gains for weaker backbones. On the post-cutoff OSS-Fuzz dataset, SWE-agent with Claude Sonnet 4 identifies 95.5% of FPs while maintaining 95.5% precision, compared with a 36.4% FP identification rate for vanilla prompting. Moreover, aggressive FP reduction can come at the cost of suppressing true vulnerabilities, highlighting important trade-offs. Finally, we observe large disparities in computational cost across agent frameworks.
Overall, our study demonstrates that LLM-based agents are a powerful but non-uniform solution for SAST FP filtering, and that their practical deployment requires careful consideration of agent design, backbone model choice, vulnerability category, and operational cost.
CCS Concepts: • Security and privacy → Software and application security.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9580a688-05a6-4a08-98dc-8faffb8bf423Builds on18
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- A Large-Scale Empirical Study of Security PatchesFrank Li, Vern PaxsonCCS 2017 · 273 citations
- RepairAgent: An Autonomous, LLM-Based Agent for Program RepairIslem Bouzenia, Premkumar T. Devanbu, Michael PradelICSE 2025 · 54 citations
- Comparison and Evaluation on Static Application Security Testing (SAST) Tools for JavaKaixuan Li, Sen Chen, Lingling Fan, Ruitao Feng et al.FSE 2023 · 43 citations
- Detecting False Alarms from Automatic Static Analysis Tools: How Far are We?Hong Jin Kang, Khai Loong Aw, David LoICSE 2022 · 42 citations
Related papers
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- Make Agent Defeat Agent: Automatic Detection of Taint-Style Vulnerabilities in LLM-based AgentsFengyu Liu, Yuan Zhang, Jiaqi Luo, Jiarun Dai et al.USENIX Security 2025
- SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing ScenariosJunkai Chen, Huihui Huang, Yunbo Lyu, Junwen An et al.ACL 2026 · 5 citations
- RefAgent: A Multi-agent LLM-based Framework for Automatic Software RefactoringKhouloud Oueslati, Maxime Lamothe, Foutse KhomhICSE 2026 · 1 citation
- When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?Yibo Peng, James Song, Lei Li, Xinyu Yang et al.ACL 2026 · 2 citations
