The Illusion of Success: Learning-Based Android Malware Detectors (Replicability Study)
Michael Tegegn, Julia Rubin
Abstract
Since 2012, hundreds of machine-learning-based classification approaches have been proposed to help separate malware from benign Android applications. These approaches typically collect a large number of applications of both types, split them into training and testing subsets, train a binary classifier on the training subset, and measure accuracy on the testing subset. They typically report very high achieved accuracy, i.e., F1-score of around 95%. Recent work has also highlighted several biases and flaws in the experimental setup and evaluation methodology of such approaches, questioning the trustworthiness of their reported results. In an effort to better understand the current status of classification-based malware detection, we first conduct a systematic literature review to extract the properties of existing tools and the datasets that they use. We then design a large-scale longitudinal study, where we evaluate the most prominent tools on a range of datasets spanning 13 years (2011-2023), which we systematically collected from the AndroZoo and VirusShare repositories while controlling for the known experimental setup biases. Our results show lower than reported classification performance, with the F1-score for a tool ranging from high 60 to low 90 percent, depending on a dataset. Moreover, even successful classification often utilizes hidden but semantically weak correlations in the data. In fact, a deliberately naïve and unreliable classifier we designed for this study, which uses application package names as features for classification, performs comparably to and sometimes even better than the state-of-the-art tools when compared under the same setup. These results challenge the claimed detection capabilities of the tools, render comparisons of reported tool accuracy (without empirical evaluation on the exact same set of applications) practically meaningless, and call for the creation of more reliable and explainable semantic malware detection tools.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5d7cc009-86df-4050-b9d9-14141540960bRelated papers
- MalWhiteout: Reducing Label Errors in Android Malware DetectionLiu Wang, Haoyu Wang, Xiapu Luo, Yulei SuiASE 2022 · 16 citations
- Continuous Learning for Android Malware DetectionYizheng Chen, Zhoujie Ding, David A. WagnerUSENIX Security 2023
- LAMDA: A Longitudinal Android Malware Benchmark for Concept Drift AnalysisMd Ahsanul Haque, Ismail Hossain, Md Mahmuduzzaman Kamol, Md Jahangir Alam et al.ICLR 2026 · 14 citations
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and TimeFeargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder et al.USENIX Security 2019 · 441 citations
- FeatureSmith: Automatically Engineering Features for Malware Detection by Mining the Security LiteratureZiyun Zhu, Tudor DumitrasCCS 2016 · 114 citations
