Lune

ISSTA2026Top-tier venue

Answer Is Cheap, Show Me the Evidence! Augmenting Automated Vulnerability Assessment with Evidence

Shengyi Pan, Zelong Zheng, Jiayuan Zhou, Xing Hu, Xin Xia, Shanping Li

2026Year

Abstract

Software Vulnerability (SV) assessment is a vital phase in SV management, which characterizes discovered SVs to locate hot spots and prioritize their remediation. To reduce the overhead and latency of manual assessment, prior works have explored automatically predicting assessment results from SV reports (SVRs). However, existing approaches fail to process the information conveyed by the rich text content (e.g., screenshots and code snippets) embedded in SVRs and miss information about vulnerable projects. More importantly, they primarily focus on assessment accuracy while neglecting to provide explanations or evidence supporting their predictions. As a result, these approaches remain impractical in real-world settings, where imperfect accuracy necessitates manual validation. LLMs offer a promising opportunity to address this limitation by performing SV assessment while simultaneously providing supporting evidence. Nevertheless, our extensive evaluation reveals that mainstream LLMs perform poorly on SV assessment tasks, largely due to a lack of assessment specific knowledge. To address the above challenges, we propose EAVA, a novel framework that effectively leverages LLMs to perform SV assessment and provide supporting evidence. EAVA employs specialized LLM agents to process rich text content in SVRs and incorporate information about vulnerable projects. EAVA builds a dedicated assessment LLM by injecting assessment-specific knowledge through finetuning. Specifically, we enable large-scale reasoning trajectory annotation using off-the-shelf LLMs and adopt a two-stage training paradigm, i.e., supervised instruction tuning to inject domain knowledge, followed by reinforcement learning to enhance the model’s intrinsic reasoning capability. Evaluations on a newly collected SVR dataset demonstrate that EAVA outperforms the best-performing baseline by 5.3%-35.2% across multiple evaluation metrics. Ablation studies validate the effectiveness of our design choices for both assessment specific model training and SV information enrichment. Finally, a user study with security experts confirms that the evidence provided by EAVA is useful and practical for real-world SV assessment.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9c2db8d4-5198-4fd6-a9ba-6c6c6fde39d6

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines