MalWhiteout: Reducing Label Errors in Android Malware Detection
Liu Wang, Haoyu Wang, Xiapu Luo, Yulei Sui
摘要
Machine learning based Android malware detection has attracted a great deal of research work in recent years. A reliable malware dataset is critical to evaluate the effectiveness of malware detection approaches. Unfortunately, existing malware datasets used in our community are mainly labelled by leveraging existing anti-virus services (i.e., VirusTotal), which are prone to mislabelling. This, however, would lead to the inaccurate evaluation of the malware detection techniques. Removing label noises from Android malware datasets can be quite challenging, especially at a large data scale. To address this problem, we propose an effective approach called MalWhiteout to reduce label errors in Android malware datasets. Specifically, we creatively introduce Confident Learning (CL), an advanced noise estimation approach, to the domain of Android malware detection. To combat false positives introduced by CL, we incorporate the idea of ensemble learning and inter-app relation to achieve a more robust capability in noise detection. We evaluate MalWhiteout on a curated large-scale and reliable benchmark dataset. Experimental results show that MalWhiteout is capable of detecting label noises with over 94% accuracy even at a high noise ratio (i.e., 30%) of the dataset. MalWhiteout outperforms the state-of-the-art approach in terms of both effectiveness (8% to 218% improvement) and efficiency (70 to 249 times faster) across different settings. By reducing label noises, we show that the performance of existing malware detection approaches can be improved.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Uncovering and Mitigating the Impact of Code Obfuscation on Dataset Annotation with Antivirus EnginesCuiying Gao, Yueming Wu, Heng Li, Wei Yuan 等ISSTA 2024 · 被引用 4 次
- Deep Learning from Imperfectly Labeled Malware DataFahad Alotaibi, Euan Goodbrand, Sergio MaffeisCCS 2025
- Understanding Model Weaknesses: A Path to Strengthening DNN-Based Android Malware DetectionHaodong Li, Xiao Cheng, Yanjie Zhao, Guosheng Xu 等ISSTA 2025
- Fighting Fire with Fire: Continuous Attack for Adversarial Android Malware DetectionYinyuan Zhang, Cuiying Gao, Yueming Wu, Shihan Dou 等USENIX Security 2025
它引用的顶会 Paper7
- MaMaDroid: Detecting Android Malware by Building Markov Chains of Behavioral ModelsEnrico Mariconti, Lucky Onwuzurike, Panagiotis Andriotis, Emiliano De Cristofaro 等NDSS 2017 · 被引用 471 次
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and TimeFeargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder 等USENIX Security 2019 · 被引用 441 次
- Can gradient clipping mitigate label noise?Aditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Sanjiv KumarICLR 2020 · 被引用 163 次
- Investigating Commercial Pay-Per-Install and the Distribution of Unwanted SoftwareKurt Thomas, Juan A. Elices Crespo, Ryan Rasti, Jean-Michel Picod 等USENIX Security 2016 · 被引用 77 次
- Towards Attribution in Mobile Markets: Identifying Developer Account PolymorphismSilvia Sebastián, Juan CaballeroCCS 2020 · 被引用 19 次
相关 Paper
- Mitigating Emergent Malware Label Noise in DNN-Based Android Malware DetectionHaodong Li, Xiao Cheng, Guohan Zhang, Guosheng Xu 等FSE 2025 · 被引用 2 次
- Differential Training: A Generic Framework to Reduce Label Noises for Android Malware DetectionJiayun Xu, Yingjiu Li, Robert H. DengNDSS 2021
- The Illusion of Success: Learning-Based Android Malware Detectors (Replicability Study)Michael Tegegn, Julia RubinISSTA 2026
- Continuous Learning for Android Malware DetectionYizheng Chen, Zhoujie Ding, David A. WagnerUSENIX Security 2023
- MalCertain: Enhancing Deep Neural Network Based Android Malware Detection by Tackling Prediction UncertaintyHaodong Li, Guosheng Xu, Liu Wang, Xusheng Xiao 等ICSE 2024 · 被引用 17 次
