Lune

ISSTA2026顶会

The Illusion of Success: Learning-Based Android Malware Detectors (Replicability Study)

Michael Tegegn, Julia Rubin

2026年份

摘要

Since 2012, hundreds of machine-learning-based classification approaches have been proposed to help separate malware from benign Android applications. These approaches typically collect a large number of applications of both types, split them into training and testing subsets, train a binary classifier on the training subset, and measure accuracy on the testing subset. They typically report very high achieved accuracy, i.e., F1-score of around 95%. Recent work has also highlighted several biases and flaws in the experimental setup and evaluation methodology of such approaches, questioning the trustworthiness of their reported results. In an effort to better understand the current status of classification-based malware detection, we first conduct a systematic literature review to extract the properties of existing tools and the datasets that they use. We then design a large-scale longitudinal study, where we evaluate the most prominent tools on a range of datasets spanning 13 years (2011-2023), which we systematically collected from the AndroZoo and VirusShare repositories while controlling for the known experimental setup biases. Our results show lower than reported classification performance, with the F1-score for a tool ranging from high 60 to low 90 percent, depending on a dataset. Moreover, even successful classification often utilizes hidden but semantically weak correlations in the data. In fact, a deliberately naïve and unreliable classifier we designed for this study, which uses application package names as features for classification, performs comparably to and sometimes even better than the state-of-the-art tools when compared under the same setup. These results challenge the claimed detection capabilities of the tools, render comparisons of reported tool accuracy (without empirical evaluation on the exact same set of applications) practically meaningless, and call for the creation of more reliable and explainable semantic malware detection tools.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 5d7cc009-86df-4050-b9d9-14141540960b

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖