Differential Training: A Generic Framework to Reduce Label Noises for Android Malware Detection
Jiayun Xu, Yingjiu Li, Robert H. Deng
摘要
—A common problem in machine learning-based malware detection is that training data may contain noisy labels and it is challenging to make the training data noise-free at a large scale. To address this problem, we propose a generic framework to reduce the noise level of training data for the training of any machine learning-based Android malware detection. Our framework makes use of all intermediate states of two identical deep learning classification models during their training with a given noisy training dataset and generate a noise-detection feature vector for each input sample. Our framework then applies a set of outlier detection algorithms on all noise-detection feature vectors to reduce the noise level of the given training data before feeding it to any machine learning based Android malware detection approach. In our experiments with three different Android malware detection approaches, our framework can detect significant portions of wrong labels in different training datasets at different noise ratios, and improve the performance of Android malware detection approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- A lightweight framework for function name reassignment based on large-scale stripped binariesHan Gao, Shaoyin Cheng, Yinxing Xue, Weiming ZhangISSTA 2021 · 被引用 46 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- MalWhiteout: Reducing Label Errors in Android Malware DetectionLiu Wang, Haoyu Wang, Xiapu Luo, Yulei SuiASE 2022 · 被引用 16 次
- Learning from Limited Heterogeneous Training Data: Meta-Learning for Unsupervised Zero-Day Web Attack Detection across Web DomainsPeiyang Li, Ye Wang, Qi Li, Zhuotao Liu 等CCS 2023 · 被引用 13 次
- Uncovering and Mitigating the Impact of Code Obfuscation on Dataset Annotation with Antivirus EnginesCuiying Gao, Yueming Wu, Heng Li, Wei Yuan 等ISSTA 2024 · 被引用 4 次
它引用的顶会 Paper1
相关 Paper
- Deep Learning from Imperfectly Labeled Malware DataFahad Alotaibi, Euan Goodbrand, Sergio MaffeisCCS 2025
- Mitigating Emergent Malware Label Noise in DNN-Based Android Malware DetectionHaodong Li, Xiao Cheng, Guohan Zhang, Guosheng Xu 等FSE 2025 · 被引用 2 次
- Guided Retraining to Enhance the Detection of Difficult Android MalwareNadia Daoudi, Kevin Allix, Tegawendé F. Bissyandé, Jacques KleinISSTA 2023 · 被引用 4 次
- MalCertain: Enhancing Deep Neural Network Based Android Malware Detection by Tackling Prediction UncertaintyHaodong Li, Guosheng Xu, Liu Wang, Xusheng Xiao 等ICSE 2024 · 被引用 17 次
- Enhancing State-of-the-art Classifiers with API Semantics to Detect Evolved Android MalwareXiaohan Zhang, Yuan Zhang, Ming Zhong, Daizong Ding 等CCS 2020 · 被引用 173 次
