An Empirical Study on the Effects of Obfuscation on Static Machine Learning-Based Malicious JavaScript Detectors
Kunlun Ren, Weizhong Qiang, Yueming Wu, Yi Zhou, Deqing Zou, Hai Jin
摘要
Machine learning is increasingly being applied to malicious JavaScript detection in response to the growing number of Web attacks and the attendant costly manual identification. In practice, to hide their malicious behaviors or protect intellectual copyrights, both malicious and benign scripts tend to obfuscate their own code before uploading. While obfuscation is beneficial, it also introduces some additional code features (e.g., dead code) into the code. When machine learning is employed to learn a malicious JavaScript detector, these additional features can affect the model to make it less effective. However, there is still a lack of clear understanding of how robust existing machine learning-based detectors are on different obfuscators. In this paper, we conduct the first empirical study to figure out how obfuscation affects machine learning detectors based on static features. Through the results, we observe several findings: 1) Obfuscation has a significant impact on the effectiveness of detectors, causing an increase both in false negative rate (FNR) and false positive rate (FPR), and the bias of obfuscation in the training set induces detectors to detect obfuscation rather than malicious behaviors. 2) The common measures such as improving the quality of the training set by adding relevant obfuscated samples and leveraging state-of-the-art deep learning models can not work well.3) The root cause of obfuscation effects on these detectors is that feature spaces they use can only reflect shallow differences in code, not about the nature of benign and malicious, which can be easily affected by the differences brought by obfuscation. 4) Obfuscation has a similar effect on realistic detectors in VirusTotal, indicating that this is a common real-world problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DONAPI: Malicious NPM Packages Detector using Behavior Sequence Knowledge MappingCheng Huang, Nannan Wang, Ziyan Wang, Siqi Sun 等USENIX Security 2024 · 被引用 38 次
- AdFlush: A Real-World Deployable Machine Learning Solution for Effective Advertisement and Web Tracker PreventionKiho Lee, Chaejin Lim, Beomjin Jin, Taeyoung Kim 等WWW 2024 · 被引用 5 次
- SpiderScan: Practical Detection of Malicious NPM Packages Based on Graph-Based Behavior Modeling and MatchingYiheng Huang, Ruisi Wang, Wen Zheng, Zhuotong Zhou 等ASE 2024 · 被引用 4 次
- Blocking Tracking JavaScript at the Function GranularityAbdul Haddi Amjad, Shaoor Munir, Zubair Shafiq, Muhammad Ali GulzarCCS 2024 · 被引用 3 次
- From Obfuscated to Obvious: A Comprehensive JavaScript Deobfuscation Tool for Security AnalysisDongchao Zhou, Lingyun Ying, Huajun Chai, Dongbin WangNDSS 2026 · 被引用 3 次
它引用的顶会 Paper8
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against ManipulationVictor Le Pochat, Tom van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczynski 等NDSS 2019 · 被引用 826 次
- Online Tracking: A 1-million-site Measurement and AnalysisSteven Englehardt, Arvind NarayananCCS 2016 · 被引用 798 次
- MineSweeper: An In-depth Look into Drive-by Cryptocurrency Mining and Its DefenseRadhesh Krishnan Konoth, Emanuele Vineti, Veelasha Moonsamy, Martina Lindorfer 等CCS 2018 · 被引用 162 次
- HideNoSeek: Camouflaging Malicious JavaScript in Benign ASTsAurore Fass, Michael Backes, Ben StockCCS 2019 · 被引用 78 次
相关 Paper
- Uncovering and Mitigating the Impact of Code Obfuscation on Dataset Annotation with Antivirus EnginesCuiying Gao, Yueming Wu, Heng Li, Wei Yuan 等ISSTA 2024 · 被引用 4 次
- When Malware is Packin' Heat; Limits of Machine Learning Classifiers Based on Static Analysis FeaturesHojjat Aghakhani, Fabio Gritti, Francesco Mecca, Martina Lindorfer 等NDSS 2020
- Wobfuscator: Obfuscating JavaScript Malware via Opportunistic Translation to WebAssemblyAlan Romano, Daniel Lehmann, Michael Pradel, Weihang WangS&P 2022 · 被引用 40 次
- Extract Me If You Can: Abusing PDF Parsers in Malware DetectorsCurtis Carmony, Xunchao Hu, Heng Yin, Abhishek Vasisht Bhaskar 等NDSS 2016 · 被引用 61 次
- ProfMal: Detecting Malicious NPM Packages by the Synergy between Static and Dynamic AnalysisYiheng Huang, Wen Zheng, Susheng Wu, Bihuan Chen 等ASE 2025 · 被引用 2 次
