When Malware is Packin' Heat; Limits of Machine Learning Classifiers Based on Static Analysis Features
Hojjat Aghakhani, Fabio Gritti, Francesco Mecca, Martina Lindorfer, Stefano Ortolani, Davide Balzarotti, Giovanni Vigna, Christopher Kruegel
摘要
Machine learning techniques are widely used in addition to signatures and heuristics to increase the detection rate of anti-malware software, as they automate the creation of detection models, making it possible to handle an ever-increasing number of new malware samples. In order to foil the analysis of anti-malware systems and evade detection, malware uses packing and other forms of obfuscation. However, few realize that benign applications use packing and obfuscation as well, to protect intellectual property and prevent license abuse.
In this paper, we study how machine learning based on static analysis features operates on packed samples. Malware researchers have often assumed that packing would prevent machine learning techniques from building effective classifiers. However, both industry and academia have published results that show that machine-learning-based classifiers can achieve good detection rates, leading many experts to think that classifiers are simply detecting the fact that a sample is packed, as packing is more prevalent in malicious samples. We show that, different from what is commonly assumed, packers do preserve some information when packing programs that is “useful” for malware classification. However, this information does not necessarily capture the sample’s behavior. We demonstrate that the signals extracted from packed executables are not rich enough for machine-learning-based models to (1) generalize their knowledge to operate on unseen packers, and (2) be robust against adversarial examples. We also show that a naïve application of machine learning techniques results in a substantial number of false positives, which, in turn, might have resulted in incorrect labeling of ground-truth data used in past work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi 等USENIX Security 2021 · 被引用 241 次
- DEEPCASE: Semi-Supervised Contextual Analysis of Security EventsThijs van Ede, Hojjat Aghakhani, Noah Spahn, Riccardo Bortolameotti 等S&P 2022 · 被引用 91 次
- Survivalism: Systematic Analysis of Windows Malware Living-Off-The-LandFrederick Barr-Smith, Xabier Ugarte-Pedrero, Mariano Graziano, Riccardo Spolaor 等S&P 2021 · 被引用 73 次
- MalGraph: Hierarchical Graph Neural Networks for Robust Windows Malware DetectionXiang Ling, Lingfei Wu, Wei Deng, Zhenqing Qu 等INFOCOM 2022 · 被引用 47 次
- How Did That Get In My Phone? Unwanted App Distribution on Android DevicesPlaton Kotzias, Juan Caballero, Leyla BilgeS&P 2021 · 被引用 37 次
它引用的顶会 Paper7
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and TimeFeargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder 等USENIX Security 2019 · 被引用 441 次
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su 等CCS 2018 · 被引用 336 次
- Transcend: Detecting Concept Drift in Malware Classification ModelsRoberto Jordaney, Kumar Sharad, Santanu Kumar Dash, Zhi Wang 等USENIX Security 2017 · 被引用 325 次
- BinSim: Trace-based Semantic Binary Diffing via System Call Sliced Segment Equivalence CheckingJiang Ming, Dongpeng Xu, Yufei Jiang, Dinghao WuUSENIX Security 2017 · 被引用 118 次
- Things You May Not Know About Android (Un)Packers: A Systematic Study based on Whole-System EmulationYue Duan, Mu Zhang, Abhishek Vasisht Bhaskar, Heng Yin 等NDSS 2018 · 被引用 87 次
相关 Paper
- Prevalence and Impact of Low-Entropy Packing Schemes in the Malware EcosystemAlessandro Mantovani, Simone Aonzo, Xabier Ugarte-Pedrero, Alessio Merlo 等NDSS 2020
- An Empirical Study on the Effects of Obfuscation on Static Machine Learning-Based Malicious JavaScript DetectorsKunlun Ren, Weizhong Qiang, Yueming Wu, Yi Zhou 等ISSTA 2023 · 被引用 11 次
- Adversarially Robust Assembly Language Model for Packed Executables DetectionShijia Li, Jiang Ming, Lanqing Liu, Longwei Yang 等CCS 2025
- Decoding the Secrets of Machine Learning in Malware Classification: A Deep Dive into Datasets, Feature Extraction, and Model PerformanceSavino Dambra, Yufei Han, Simone Aonzo, Platon Kotzias 等CCS 2023 · 被引用 28 次
- Adversarial Training for Raw-Binary Malware ClassifiersKeane Lucas, Samruddhi Pai, Weiran Lin, Lujo Bauer 等USENIX Security 2023
