Understanding Model Weaknesses: A Path to Strengthening DNN-Based Android Malware Detection
Haodong Li, Xiao Cheng, Yanjie Zhao, Guosheng Xu, Guoai Xu, Haoyu Wang
摘要
Android malware detection remains a critical challenge in cybersecurity research. Recent advancements leverage AI techniques, particularly deep neural networks (DNNs), to train a detection model, but their effectiveness is often compromised by the pronounced imbalance among malware families in commonly used training datasets. This imbalance leads to overfitting in dominant categories and poor performance in underrepresented ones, increasing predictive uncertainty for less common malware families. To address the suboptimal performance of many DNN models, we introduce MalTutor, a novel framework that enhances model robustness through an optimized training process. Our primary insight lies in transforming uncertainties from "liabilities" into "assets" by strategically incorporating them into DNN training methodologies. Specifically, we begin by evaluating the predictive uncertainty of DNN models throughout various training epochs, which guides our sample categorization. Incorporating Curriculum Learning strategies, we commence training with easy-to-learn samples with lower uncertainty, progressively incorporating difficult-to-learn ones with higher uncertainty. Our experimental results demonstrate that MalTutor significantly improves the performance of models trained on imbalanced datasets, increasing accuracy by 31.0%, elevating the F1 score by 138.8%, and specifically boosting the average accuracy in detecting various types of malicious apps by 133.9%. Our findings provide valuable insights into the potential benefits of incorporating uncertainty to improve the robustness of DNN models for prediction-oriented software engineering tasks. CCS Concepts: • Security and privacy → Malware and its mitigation; • Computing methodologies → Learning paradigms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi 等USENIX Security 2021 · 被引用 241 次
- Curriculum Pre-training for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Ming Zhou 等ACL 2020 · 被引用 100 次
- Towards characterizing adversarial defects of deep learning software from the lens of uncertaintyXiyue Zhang, Xiaofei Xie, Lei Ma, Xiaoning Du 等ICSE 2020 · 被引用 69 次
- Out of Distribution Data Detection Using Dropout Bayesian Neural NetworksAndré T. Nguyen, Fred Lu, Gary Lopez Munoz, Edward Raff 等AAAI 2022 · 被引用 30 次
- Does data sampling improve deep learning-based vulnerability detection? Yeas! and Nays!Xu Yang, Shaowei Wang, Yi Li, Shaohua WangICSE 2023 · 被引用 19 次
相关 Paper
- MalCertain: Enhancing Deep Neural Network Based Android Malware Detection by Tackling Prediction UncertaintyHaodong Li, Guosheng Xu, Liu Wang, Xusheng Xiao 等ICSE 2024 · 被引用 17 次
- Mitigating Emergent Malware Label Noise in DNN-Based Android Malware DetectionHaodong Li, Xiao Cheng, Guohan Zhang, Guosheng Xu 等FSE 2025 · 被引用 2 次
- Adversarial Training for Raw-Binary Malware ClassifiersKeane Lucas, Samruddhi Pai, Weiran Lin, Lujo Bauer 等USENIX Security 2023
- Deep Learning from Imperfectly Labeled Malware DataFahad Alotaibi, Euan Goodbrand, Sergio MaffeisCCS 2025
- Guided Retraining to Enhance the Detection of Difficult Android MalwareNadia Daoudi, Kevin Allix, Tegawendé F. Bissyandé, Jacques KleinISSTA 2023 · 被引用 4 次
