When a Tree Falls: Using Diversity in Ensemble Classifiers to Identify Evasion in Malware Detectors
Charles Smutz, Angelos Stavrou
Abstract
Machine learning classifiers are a vital component of modern malware and intrusion detection systems. However, past studies have shown that classifier based detection systems are susceptible to evasion attacks in practice. Improving the evasion resistance of learning based systems is an open problem. To address this, we introduce a novel method for identifying the observations on which an ensemble classifier performs poorly. During detection, when a sufficient number of votes from individual classifiers disagree, the ensemble classifier prediction is shown to be unreliable. The proposed method, ensemble classifier mutual agreement analysis, allows the detection of many forms of classifier evasion without additional external ground truth. We evaluate our approach using PDFrate, a PDF malware detector. Applying our method to data taken from a real network, we show that the vast majority of predictions can be made with high ensemble classifier agreement. However, most classifier evasion attempts, including nine targeted mimicry scenarios from two recent studies, are given an outcome of uncertain indicating that these observations cannot be given a reliable prediction by the classifier. To show the general applicability of our approach, we tested it against the Drebin Android malware detector where an uncertain prediction was correctly given to the majority of novel attacks. Our evaluation includes over 100,000 PDF documents and 100,000 Android applications. Furthermore, we show that our approach can be generalized to weaken the effectiveness of the Gradient Descent and Kernel Density Estimation attacks against Support Vector Machines. We discovered that feature bagging is the most important property for enabling ensemble classifier diversity based evasion detection. Permission to freely reproduce all or part of this paper for noncommercial purposes is granted provided that copies bear this notice and the full citation on the first page. Reproduction for commercial purposes is strictly prohibited without the prior written consent of the Internet Society, the first-named author (for reproduction of an entire paper only), and the author's employer if the paper was prepared within the scope of employment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 959e4e1e-c1b7-4530-a9b4-5608f7fe3c9fCited by top-tier papers8
- HideNoSeek: Camouflaging Malicious JavaScript in Benign ASTsAurore Fass, Michael Backes, Ben StockCCS 2019 · 78 citations
- Deep Partition Aggregation: Provable Defenses against General Poisoning AttacksAlexander Levine, Soheil FeiziICLR 2021 · 22 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Density Boosts Everything: A One-stop Strategy for Improving Performance, Robustness, and Sustainability of Malware DetectorsJianwen Tian, Wei Kong, Debin Gao, Tong Wang et al.NDSS 2025
- Fighting Fire with Fire: Continuous Attack for Adversarial Android Malware DetectionYinyuan Zhang, Cuiying Gao, Yueming Wu, Shihan Dou et al.USENIX Security 2025
Related papers
- Automatically Evading Classifiers: A Case Study on PDF Malware ClassifiersWeilin Xu, Yanjun Qi, David EvansNDSS 2016 · 249 citations
- Improving Robustness of ML Classifiers against Realizable Evasion Attacks Using Conserved FeaturesLiang Tong, Bo Li, Chen Hajaj, Chaowei Xiao et al.USENIX Security 2019 · 95 citations
- Evading Classifiers by Morphing in the DarkHung Dang, Yue Huang, Ee-Chien ChangCCS 2017 · 125 citations
- Automated Mass Malware Factory: The Convergence of Piggybacking and Adversarial Example in Android Malicious Software GenerationHeng Li, Zhiyuan Yao, Bang Wu, Cuiying Gao et al.NDSS 2025
- DRSM: De-Randomized Smoothing on Malware Classifier Providing Certified RobustnessShoumik Saha, Wenxiao Wang, Yigitcan Kaya, Soheil Feizi et al.ICLR 2024 · 6 citations
