Argus: A Multi-agent Sensitive Information Leakage Detection Framework Based on Hierarchical Reference Relationships
Wang Bin, Hui Li, Liyang Zhang, Qijia Zhuang, Ao Yang, Dong Zhang, Xijun Luo, Bing Lin
摘要
Sensitive information leakage in code repositories has emerged as a critical security challenge. Traditional detection methods—relying on regular expressions, fingerprint features, and high-entropy calculations-suffer from high false-positive rates, which not only reduce detection efficiency but also significantly increase the manual screening burden on developers. Recent advances in large language models(LLMs) and multi-agent collaborative architectures have demonstrated remarkable potential in tackling complex tasks, offering a novel technological perspective for sensitive information detection. In response to these challenges, we propose Argus, a multi-agent collaborative framework for detecting sensitive information. Argus employs a three-tier detection mechanism that integrates key content, file context, and project reference relationships to effectively reduce false positives and enhance overall detection accuracy. To comprehensively evaluate Argus in real-world repository environments, we developed two new benchmarks—one to assess genuine leak detection capabilities and another to evaluate false-positive filtering performance. Experimental results show that Argus achieves up to 94.86% accuracy in leak detection, with a precision of 96.36%, recall of 94.64%, and an F1 score of 0.955.Moreover, the analysis of 97 real repositories incurred a total cost of only $2.21. All code implementations and related datasets are publicly available at https://github.com/TheBinKing/Argus-Guard for further research and application.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- How Bad Can It Git? Characterizing Secret Leakage in Public GitHub RepositoriesMichael Meli, Matthew R. McNiece, Bradley ReavesNDSS 2019 · 被引用 130 次
- Teach AI How to Code: Using Large Language Models as Teachable Agents for Programming EducationHyoungwook Jin, Seonghee Lee, Hyungyu Shin, Juho KimCHI 2024 · 被引用 94 次
- Detecting False Alarms from Automatic Static Analysis Tools: How Far are We?Hong Jin Kang, Khai Loong Aw, David LoICSE 2022 · 被引用 42 次
- "False negative - that one is going to kill you": Understanding Industry Perspectives of Static Analysis based Security TestingAmit Seal Ami, Kevin Moran, Denys Poshyvanyk, Adwait NadkarniS&P 2024 · 被引用 40 次
- Automated Detection of Password Leakage from Public GitHub RepositoriesRunhan Feng, Ziyang Yan, Shiyan Peng, Yuanyuan ZhangICSE 2022 · 被引用 36 次
相关 Paper
- Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability DetectionXin Peng, Bo Lin, Jing Wang, Xiaoling Li 等FSE 2026 · 被引用 1 次
- Goal-Aware Identification and Rectification of Misinformation in Multi-Agent SystemsZherui Li, Yan Mi, Zhenhong Zhou, Houcheng Jiang 等ICLR 2026 · 被引用 10 次
- Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code RepositoriesAlperen Yildiz, Sin G. Teo, Yiling Lou, Yebo Feng 等ACL 2025 · 被引用 32 次
- RepoAudit: An Autonomous LLM-Agent for Repository-Level Code AuditingJinyao Guo, Chengpeng Wang, Xiangzhe Xu, Zian Su 等ICML 2025
- Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive FilteringYunpeng Xiong, Ting ZhangISSTA 2026
