Uncovering and Mitigating the Impact of Code Obfuscation on Dataset Annotation with Antivirus Engines
Cuiying Gao, Yueming Wu, Heng Li, Wei Yuan, Haoyu Jiang, Qidan He, Yang Liu
摘要
With the widespread application of machine learning-based Android malware detection methods, building a high-quality dataset has become increasingly important. Existing large-scale datasets are mostly annotated with VirusTotal by aggregating the decisions of antivirus engines, and most of them indiscriminately accept the decisions of all engines. In reality, however, these engines have different capabilities in detecting malware, especially those that have been obfuscated. Previous research has revealed that code obfuscation degrades the detection performance of these engines to varying degrees. This makes us believe that using all engines indiscriminately is unreasonable for dataset annotation. Therefore, in this paper, we first conduct a data-driven evaluation to confirm the negative effects of code obfuscation on engine-based dataset annotation. To gain a deeper understanding of the reasons behind this phenomenon, we evaluate the availability, effectiveness and robustness of every engine under various code obfuscation techniques. Then we categorize the engines and select a set of obfuscation-robust engines. Finally, we conduct comprehensive experiments to verify the effectiveness of the selected engines for dataset annotation. Our experiments show that when 50% obfuscated samples are mixed into the training set, on the classic malware detectors Drebin and Malscan, using our selected engines can effectively improve detection performance by 15.21% and 19.23%, respectively, compared to using all the engines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- MaMaDroid: Detecting Android Malware by Building Markov Chains of Behavioral ModelsEnrico Mariconti, Lucky Onwuzurike, Panagiotis Andriotis, Emiliano De Cristofaro 等NDSS 2017 · 被引用 471 次
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and TimeFeargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder 等USENIX Security 2019 · 被引用 441 次
- Intriguing Properties of Adversarial ML Attacks in the Problem SpaceFabio Pierazzi, Feargus Pendlebury, Jacopo Cortellazzi, Lorenzo CavallaroS&P 2020 · 被引用 334 次
- 50 Ways to Leak Your Data: An Exploration of Apps' Circumvention of the Android Permissions SystemJoel Reardon, Álvaro Feal, Primal Wijesekera, Amit Elazari Bar On 等USENIX Security 2019 · 被引用 196 次
- Unleashing the hidden power of compiler optimization on binary code difference: an empirical studyXiaolei Ren, Michael Ho, Jiang Ming, Yu Lei 等PLDI 2021 · 被引用 57 次
相关 Paper
- An Empirical Study on the Effects of Obfuscation on Static Machine Learning-Based Malicious JavaScript DetectorsKunlun Ren, Weizhong Qiang, Yueming Wu, Yi Zhou 等ISSTA 2023 · 被引用 11 次
- Measuring and Modeling the Label Dynamics of Online Anti-Malware EnginesShuofei Zhu, Jianjun Shi, Limin Yang, Boqin Qin 等USENIX Security 2020
- A Comprehensive Study of Learning-based Android Malware Detectors under Challenging EnvironmentsCuiying Gao, Gaozhun Huang, Heng Li, Bang Wu 等ICSE 2024 · 被引用 29 次
- MalWhiteout: Reducing Label Errors in Android Malware DetectionLiu Wang, Haoyu Wang, Xiapu Luo, Yulei SuiASE 2022 · 被引用 16 次
- Tackling runtime-based obfuscation in Android with TIROMichelle Y. Wong, David LieUSENIX Security 2018 · 被引用 59 次
