Is “Knowing It’s Malicious” Enough? Evaluating LLMs for Fine-Grained Malware Behavior Auditing
Xinran Zheng, Xingzhi Qian, Yiling He, Shuo Yang, Lorenzo Cavallaro
Abstract
Automated malware classifiers achieve strong detection performance, but auditing requires more than flagging a sample: analysts must explain malicious behaviors and justify them with concrete code evidence, a requirement that traditional signature-based methods and learning-based XAI often fail to satisfy in a human-interpretable manner. Large Language Models (LLMs) appear well-suited for this task due to their code reasoning and summarization ability, yet it remains unclear whether they can support reliable auditing. In particular, evaluating them faces three hurdles: (1) the lack of detailed, human-written behavior ground truth for reliable benchmarking; (2) real-world application codebases typically exceed the context limits of current models, which cannot be fully processed at once; and (3) the absence of reliable mechanisms to verify whether LLM-generated behavioral claims are faithfully supported by concrete code evidence.
Together, these obstacles make benchmarking LLM-based auditing non-trivial, leaving their true capabilities and failure modes opaque. To bridge this gap, we introduce MalEval, a diagnostic evaluation framework for systematically measuring the capability boundaries of LLMs in malware auditing. We pair real-world application codebases with expert-written audit reports to obtain fine-grained, behavior-level ground truth. Large codebases are compressed into unified behavior-relevant program contexts via a context-driven intermediate representation that preserves essential call relations. Both expert reports and model outputs are then mapped, through constrained reasoning, into structured evidence chains linking code-level facts to high-level behaviors in a common, comparable space. Built on this foundation, MalEval decomposes auditing into 4 stage-wise auditing tasks, allowing each intermediate judgment to be independently verified under limited context windows. We leverage MalEval to evaluate seven widely used LLMs and uncover clear capability boundaries: models rely on surface cues over verifiable evidence, struggle to compose dispersed facts into coherent attack chains, and are highly sensitive to context formulation. These findings shift the focus from optimizing isolated outputs to designing LLM and agentic workflows that can reliably support malware auditing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0626befb-ebe3-433c-a231-c6cd612235c6Builds on14
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su et al.CCS 2018 · 336 citations
- Transcending TRANSCEND: Revisiting Malware Classification in the Presence of Concept DriftFederico Barbero, Feargus Pendlebury, Fabio Pierazzi, Lorenzo CavallaroS&P 2022 · 124 citations
- Large Language Models for Code: Security Hardening and Adversarial TestingJingxuan He, Martin T. VechevCCS 2023 · 98 citations
Related papers
- Beyond Raw Bytes: Towards Large Malware Language ModelsLuke Kurlandski, Harel Berger, Yin Pan, Matthew WrightNDSS 2026 · 5 citations
- Large Language Models for Code Analysis: Do LLMs Really Do Their Job?Chongzhou Fang, Ning Miao, Shaurya Srivastav, Jialin Liu et al.USENIX Security 2024 · 110 citations
- RepoAudit: An Autonomous LLM-Agent for Repository-Level Code AuditingJinyao Guo, Chengpeng Wang, Xiangzhe Xu, Zian Su et al.ICML 2025
- Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented ScanningShenao Yan, Shan Jin, Shimaa Ahmed, Sunpreet Singh Arora et al.CCS 2026
- Reasoning Runtime Behavior of a Program with LLM: How Far are We?Junkai Chen, Zhiyuan Pan, Xing Hu, Zhenhao Li et al.ICSE 2025 · 5 citations
