Answer Is Cheap, Show Me the Evidence! Augmenting Automated Vulnerability Assessment with Evidence
Shengyi Pan, Zelong Zheng, Jiayuan Zhou, Xing Hu, Xin Xia, Shanping Li
摘要
Software Vulnerability (SV) assessment is a vital phase in SV management, which characterizes discovered SVs to locate hot spots and prioritize their remediation. To reduce the overhead and latency of manual assessment, prior works have explored automatically predicting assessment results from SV reports (SVRs). However, existing approaches fail to process the information conveyed by the rich text content (e.g., screenshots and code snippets) embedded in SVRs and miss information about vulnerable projects. More importantly, they primarily focus on assessment accuracy while neglecting to provide explanations or evidence supporting their predictions. As a result, these approaches remain impractical in real-world settings, where imperfect accuracy necessitates manual validation. LLMs offer a promising opportunity to address this limitation by performing SV assessment while simultaneously providing supporting evidence. Nevertheless, our extensive evaluation reveals that mainstream LLMs perform poorly on SV assessment tasks, largely due to a lack of assessment specific knowledge. To address the above challenges, we propose EAVA, a novel framework that effectively leverages LLMs to perform SV assessment and provide supporting evidence. EAVA employs specialized LLM agents to process rich text content in SVRs and incorporate information about vulnerable projects. EAVA builds a dedicated assessment LLM by injecting assessment-specific knowledge through finetuning. Specifically, we enable large-scale reasoning trajectory annotation using off-the-shelf LLMs and adopt a two-stage training paradigm, i.e., supervised instruction tuning to inject domain knowledge, followed by reinforcement learning to enhance the model’s intrinsic reasoning capability. Evaluations on a newly collected SVR dataset demonstrate that EAVA outperforms the best-performing baseline by 5.3%-35.2% across multiple evaluation metrics. Ablation studies validate the effectiveness of our design choices for both assessment specific model training and SV information enrichment. Finally, a user study with security experts confirms that the evidence provided by EAVA is useful and practical for real-world SV assessment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- DeepCVA: Automated Commit-level Vulnerability Assessment with Deep Multi-task LearningTriet Huynh Minh Le, David Hin, Roland Croft, Muhammad Ali BabarASE 2021 · 被引用 62 次
- Automated unearthing of dangerous issue reportsShengyi Pan, Jiayuan Zhou, Filipe Roseiro Côgo, Xin Xia 等FSE 2022 · 被引用 22 次
- Commit-Level, Neural Vulnerability Detection and AssessmentYi Li, Aashish Yadavally, Jiaxing Zhang, Shaohua Wang 等FSE 2023 · 被引用 12 次
- Towards More Practical Automation of Vulnerability AssessmentShengyi Pan, Lingfeng Bao, Jiayuan Zhou, Xing Hu 等ICSE 2024 · 被引用 8 次
相关 Paper
- Vul-R2: A Reasoning LLM for Automated Vulnerability RepairXin-Cheng Wen, Zirui Lin, Yijun Yang, Cuiyun Gao 等ASE 2025 · 被引用 1 次
- Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM AgentsPengfei He, Ash Fox, Lesly Miculicich, Stefan Friedli 等ICML 2026 · 被引用 10 次
- PATCHAGENT: A Practical Program Repair Agent Mimicking Human ExpertiseZheng Yu, Ziyi Guo, Yuhang Wu, Jiahao Yu 等USENIX Security 2025
- Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability DetectionXin Peng, Bo Lin, Jing Wang, Xiaoling Li 等FSE 2026 · 被引用 1 次
- RealVul: Can We Detect Vulnerabilities in Web Applications with LLM?Di Cao, Yong Liao, Xiuwei ShangEMNLP 2024 · 被引用 11 次
