EMC: A Semantic-Enhanced Malware Classification Method with Robustness and Scalability
Haojun Zhao, Yueming Wu, Zhen Li, Deqing Zou
Abstract
Driven by substantial financial incentives, Portable Executable (PE) malware continues to evolve, posing persistent and growing threats. Malware classification has long been a vital research area, and numerous classification approaches have been proposed. However, existing methods still face performance limitations and often lack robustness when dealing with complex analytical scenarios that involve concept drift. In this paper, we propose a robust and scalable PE malware classification framework based on program semantic analysis and feature enhancement, and implement a prototype system named EMC. Our approach focuses on behavior-oriented semantic understanding of programs, constructing a more effective feature space while eliminating spurious correlations between features at a fine-grained level to enhance robustness. EMC extracts comprehensive behavior-oriented binary opcode sequences and employs an encoder sliding window mechanism for semantic understanding and feature space construction. Furthermore, it combines random Fourier features and weighted resampling techniques to remove dependencies between features, and leverages mutual information to purify features. These enhancements enable the classification model to more accurately capture the intrinsic characteristics of malicious programs, thereby improving both accuracy and robustness. Compared with ten mainstream malware classification methods, EMC improves F1 scores by up to 12.06% under normal scenarios and by 6.64%–50.35% in scenarios involving concept drift.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3aa20ce6-807b-4681-a3b6-3f08bcb5fd9eRelated papers
- DRSM: De-Randomized Smoothing on Malware Classifier Providing Certified RobustnessShoumik Saha, Wenxiao Wang, Yigitcan Kaya, Soheil Feizi et al.ICLR 2024 · 6 citations
- Transcend: Detecting Concept Drift in Malware Classification ModelsRoberto Jordaney, Kumar Sharad, Santanu Kumar Dash, Zhi Wang et al.USENIX Security 2017 · 325 citations
- Improving Robustness of ML Classifiers against Realizable Evasion Attacks Using Conserved FeaturesLiang Tong, Bo Li, Chen Hajaj, Chaowei Xiao et al.USENIX Security 2019 · 95 citations
- Continuous Learning for Android Malware DetectionYizheng Chen, Zhoujie Ding, David A. WagnerUSENIX Security 2023
- Combating Concept Drift with Explanatory Detection and Adaptation for Android Malware ClassificationYiling He, Junchi Lei, Zhan Qin, Kui Ren et al.CCS 2025 · 2 citations
