USENIX Security2023Top-tier venue
Humans vs. Machines in Malware Classification
Simone Aonzo, Yufei Han, Alessandro Mantovani, Davide Balzarotti
Abstract
Today, the classification of a file as either benign or malicious is performed by a combination of deterministic indicators (such as antivirus rules), Machine Learning classifiers, and, more importantly, the judgment of human experts. However, to compare the difference between human and machine intelligence in malware analysis, it is first necessary to understand how human subjects approach malware classification. In this direction, our work presents the first experimental study designed to capture which 'features' of a suspicious program (e.g., static properties or runtime behaviors) are prioritized for malware classification according to humans and machines intelligence. For this purpose, we created a malware classification game where 110 human players worldwide and with different seniority levels (72 novices and 38 experts) have competed to classify the highest number of unknown samples based on detailed sandbox reports. Surprisingly, we discovered that both experts and novices base their decisions on approximately the same features, even if there are clear differences between the two expertise classes. Furthermore, we implemented two state-of-the-art Machine Learning models for malware classification and evaluated their performances on the same set of samples. The comparative analysis of the results unveiled a common set of features preferred by both Machine Learning models and helped better understand the difference in the feature extraction. This work reflects the difference in the decision-making process of humans and computer algorithms and the different ways they extract information from the same data. Its findings serve multiple purposes, from training better malware analysts to improving feature encoding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d159253-dc00-4696-899c-f3c281f4a3eaCited by top-tier papers9
- Decoding the Secrets of Machine Learning in Malware Classification: A Deep Dive into Datasets, Feature Extraction, and Model PerformanceSavino Dambra, Yufei Han, Simone Aonzo, Platon Kotzias et al.CCS 2023 · 28 citations
- ShieldedCode: Learning Robust Representations for Virtual Machine Protected CodeMingqiao Mo, Yunlong Tan, Hao Zhang, Heng Zhang et al.ICLR 2026 · 10 citations
- Decompiling the Synergy: An Empirical Study of Human-LLM Teaming in Software Reverse EngineeringZion Leonahenahe Basque, Samuele Doria, Ananta Soneji, Wil Gibbs et al.NDSS 2026 · 7 citations
- Combating Concept Drift with Explanatory Detection and Adaptation for Android Malware ClassificationYiling He, Junchi Lei, Zhan Qin, Kui Ren et al.CCS 2025 · 2 citations
- Expert Insights into Advanced Persistent Threats: Analysis, Attribution, and ChallengesAakanksha Saha, James Mattei, Jorge Blasco, Lorenzo Cavallaro et al.USENIX Security 2025
Builds on8
- Asymmetric Shapley values: incorporating causal knowledge into model-agnostic explainabilityChristopher Frye, Colin Rowat, Ilya FeigeNeurIPS 2020 · 246 citations
- Understanding Linux MalwareEmanuele Cozzi, Mariano Graziano, Yanick Fratantonio, Davide BalzarottiS&P 2018 · 203 citations
- Helping Johnny to Analyze Malware: A Usability-Optimized Decompiler and Malware Analysis User StudyKhaled Yakdan, Sergej Dechand, Elmar Gerhards-Padilla, Matthew SmithS&P 2016 · 128 citations
- An Inside Look into the Practice of Malware AnalysisMiuyin Yong Wong, Matthew Landen, Manos Antonakakis, Douglas M. Blough et al.CCS 2021 · 58 citations
- Learning to Execute Programs with Instruction Pointer Attention Graph Neural NetworksDavid Bieber, Charles Sutton, Hugo Larochelle, Daniel TarlowNeurIPS 2020 · 51 citations
Related papers
- Does Every Second Count? Time-based Evolution of Malware Behavior in SandboxesAlexander Küchler, Alessandro Mantovani, Yufei Han, Leyla Bilge et al.NDSS 2021
- The Illusion of Success: Learning-Based Android Malware Detectors (Replicability Study)Michael Tegegn, Julia RubinISSTA 2026
- "I'm regretting that I hit run": In-situ Assessment of Potential MalwareBrandon Lit, Edward Crowder, Hassan Khan, Daniel VogelUSENIX Security 2025
- When Malware is Packin' Heat; Limits of Machine Learning Classifiers Based on Static Analysis FeaturesHojjat Aghakhani, Fabio Gritti, Francesco Mecca, Martina Lindorfer et al.NDSS 2020
- Prevalence and Impact of Low-Entropy Packing Schemes in the Malware EcosystemAlessandro Mantovani, Simone Aonzo, Xabier Ugarte-Pedrero, Alessio Merlo et al.NDSS 2020
