Decoding the Secrets of Machine Learning in Malware Classification: A Deep Dive into Datasets, Feature Extraction, and Model Performance
Savino Dambra, Yufei Han, Simone Aonzo, Platon Kotzias, Antonino Vitale, Juan Caballero, Davide Balzarotti, Leyla Bilge
摘要
Many studies have proposed machine-learning (ML) models for malware detection and classification, reporting an almost-perfect performance. However, they assemble ground-truth in different ways, use diverse static-and dynamic-analysis techniques for feature extraction, and even differ on what they consider a malware family. As a consequence, our community still lacks an understanding of malware classification results: whether they are tied to the nature and distribution of the collected dataset, to what extent the number of families and samples in the training dataset influence performance, and how well static and dynamic features complement each other. This work sheds light on those open questions by investigating the impact of datasets, features, and classifiers on ML-based malware detection and classification. For this, we collect the largest balanced malware dataset so far with 67k samples from 670 families (100 samples each), and train state-of-the-art models for malware detection and family classification using our dataset. Our results reveal that static features perform better than dynamic features, and that combining both only provides marginal improvement over static features. We discover no correlation between packing and classification accuracy, and that missing behaviors in dynamically-extracted features highly penalise their performance. We also demonstrate how a larger number of families to classify makes the classification harder, while a higher number of samples per family increases accuracy. Finally, we find that models trained on a uniform distribution of samples per family better generalize on unseen data. CCS CONCEPTS • Security and privacy → Malware and its mitigation;
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Beyond Raw Bytes: Towards Large Malware Language ModelsLuke Kurlandski, Harel Berger, Yin Pan, Matthew WrightNDSS 2026 · 被引用 5 次
- Efficient Code Analysis via Graph Representation Learning-Guided Large Language ModelsHang Gao, Tao Peng, Baoquan Cui, Hong Huang 等ICML 2026 · 被引用 1 次
- KEENHash: Hashing Programs into Function-Aware Embeddings for Large-Scale Binary Code Similarity AnalysisZhijie Liu, Qiyi Tang, Sen Nie, Shi Wu 等ISSTA 2025 · 被引用 1 次
- AutoMalDesc: Large-Scale Script Analysis for Cyber Threat ResearchAlexandru-Mihai Apostu, Andrei Preda, Alexandra Daniela Damir, Diana Bolocan 等AAAI 2026
- Towards Generality: Task-Adaptive Binary Analysis via Semantic Retrieval and Verifiable ReasoningYuzhe Liu, Zhijie Liu, Zhengmin Yu, Shu Wang 等USENIX Security 2026
它引用的顶会 Paper11
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and TimeFeargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder 等USENIX Security 2019 · 被引用 441 次
- Transcend: Detecting Concept Drift in Malware Classification ModelsRoberto Jordaney, Kumar Sharad, Santanu Kumar Dash, Zhi Wang 等USENIX Security 2017 · 被引用 325 次
- Dynamic Malware Analysis with Feature Engineering and Feature LearningZhaoqi Zhang, Panpan Qi, Wei WangAAAI 2020 · 被引用 153 次
- Spotless Sandboxes: Evading Malware Analysis Systems Using Wear-and-Tear ArtifactsNajmeh Miramirkhani, Mahathi Priya Appini, Nick Nikiforakis, Michalis PolychronakisS&P 2017 · 被引用 134 次
相关 Paper
- When Malware is Packin' Heat; Limits of Machine Learning Classifiers Based on Static Analysis FeaturesHojjat Aghakhani, Fabio Gritti, Francesco Mecca, Martina Lindorfer 等NDSS 2020
- Prevalence and Impact of Low-Entropy Packing Schemes in the Malware EcosystemAlessandro Mantovani, Simone Aonzo, Xabier Ugarte-Pedrero, Alessio Merlo 等NDSS 2020
- The Illusion of Success: Learning-Based Android Malware Detectors (Replicability Study)Michael Tegegn, Julia RubinISSTA 2026
- Unsuccessful story about few shot malware family classification and siamese network to the rescueYude Bai, Zhenchang Xing, Xiaohong Li, Zhiyong Feng 等ICSE 2020 · 被引用 26 次
- LAMDA: A Longitudinal Android Malware Benchmark for Concept Drift AnalysisMd Ahsanul Haque, Ismail Hossain, Md Mahmuduzzaman Kamol, Md Jahangir Alam 等ICLR 2026 · 被引用 14 次
