USENIX Security2023Top-tier venue
Adversarial Training for Raw-Binary Malware Classifiers
Keane Lucas, Samruddhi Pai, Weiran Lin, Lujo Bauer, Michael K. Reiter, Mahmood Sharif
Abstract
Machine learning (ML) models have shown promise in classifying raw executable files (binaries) as malicious or benign with high accuracy. This has led to the increasing influence of ML-based classification methods in academic and real-world malware detection, a critical tool in cybersecurity. However, previous work provoked caution by creating variants of malicious binaries, referred to as adversarial examples, that are transformed in a functionality-preserving way to evade detection. In this work, we investigate the effectiveness of using adversarial training methods to create malware-classification models that are more robust to some state-of-the-art attacks. To train our most robust models, we significantly increase the efficiency and scale of creating adversarial examples to make adversarial training practical, which has not been done before in raw-binary malware detectors. We then analyze the effects of varying the length of adversarial training, as well as analyze the effects of training with various types of attacks. We find that data augmentation does not deter state-of-the-art attacks, but that using a generic gradient-guided method, used in other discrete domains, does improve robustness. We also show that in most cases, models can be made more robust to malware-domain attacks by adversarially training them with lower-effort versions of the same attack. In the best case, we reduce one state-of-the-art attack's success rate from 90% to 5%. We also find that training with some types of attacks can increase robustness to other types of attacks. Finally, we discuss insights gained from our results, and how they can be used to more effectively train robust malware detectors. cluding vastly increasing the number of binaries eligible to be turned into adversarial examples (200 -→ 126,009); code-level optimizations to the adversarial-example-generation code to speed up each individual attack (Sec. 3.3); a distributed system of over 140 workers across thirteen servers to produce adversarial examples in parallel (Sec. 3.1); and training with less effective but computationally cheaper versions (Sec. 3.3) of the originally proposed attacks [33, 36] . Our findings include the following: • Adversarial training using data-augmentation techniques is ineffective in defending against adversarial examples in the domain of raw-binary classification. • In contrast, adversarial training using state-of-the-art attacks can yield robust models (90% -→ 5%, 26% -→ 6%, and 84% -→ 30% attack success rate for Disp-, IPR-, Kreuk-based attacks, respectively) when evaluated against the same type of attack used during training. We also find that these models can reduce success for attacks they were not trained on by up to 65%. We investigate how changing parameters of the attacks used in training (e.g., number of attack iterations, permitted file size increase, type of attack) affects the robustness of the resulting model. • We construct, train with, and evaluate a perturbation strategy inspired by other discrete domains [61, 62] , referred to as Greedy, which modifies the discrete bytes of the input binary without regard to the underlying structure, to insert evasive bytes in the most important locations. We find that this approach is effective in inducing some robustness against all attacks. • Finally, we adversarially train a classifier with adversarial examples created by three different attacks, and find that this classifier exhibits close to best-case robustness against every attack type. Background and Related Work Our work builds on research using deep neural networks (DNNs) for malware detection from raw binaries [32, 46] , along with work generating adversarial examples in the malware domain [18, 19, 33, 36, 52 ] and other domains [5, 7, 9, 23, 41, 55] . Next, we cover the background of these topics, and describe the state of the art in creating adversarial examples for malware detectors and defending against them. DNNs for Malware Detection Various feature types have been proposed for malware detection. Expert-designed features (e.g., [3, 4, 20, 22, 26, 30, 49] ) include vectors describing imported libraries, library or API function calls, network addresses contacted, number of system calls, byte entropy, whether specific strings are present, etc. Hand-crafting these features is time-consuming, but their use in DNNs has comparable results to DNNs that learn directly from raw bytes [3, 32, 36, 46] . In this paper, we only consider DNNs that infer maliciousness directly from raw bytes. MalConv [46] and AvastNet [32] are examples of DNN architectures that use raw-bytes as input. These architectures, used in many related works [18, 21, 36, 52] , achieved 98.5% and 98.6% test accuracies respectively in discriminating unseen executables [36] . Real-world malware detection consists of an ensemble of detection mechanisms, not just including the static analysis these DNNs perform [56] . Other analyses include dynamic analyses of
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- DRSM: De-Randomized Smoothing on Malware Classifier Providing Certified RobustnessShoumik Saha, Wenxiao Wang, Yigitcan Kaya, Soheil Feizi et al.ICLR 2024 · 6 citations
- MalwareTotal: Multi-Faceted and Sequence-Aware Bypass Tactics against Static Malware DetectionShuai He, Cai Fu, Hong Hu, Jiahe Chen et al.ICSE 2024 · 3 citations
- Training Robust ML-based Raw-Binary Malware Detectors in Hours, not MonthsKeane Lucas, Weiran Lin, Lujo Bauer, Michael K. Reiter et al.CCS 2024 · 2 citations
- Fool Me If You Can: On the Robustness of Binary Code Similarity Detection Models against Semantics-Preserving TransformationsJiyong Uhm, Minseok Kim, Michalis Polychronakis, Hyungjoon KooFSE 2026 · 1 citation
- Density Boosts Everything: A One-stop Strategy for Improving Performance, Robustness, and Sustainability of Malware DetectorsJianwen Tian, Wei Kong, Debin Gao, Tong Wang et al.NDSS 2025
Builds on6
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Intriguing Properties of Adversarial ML Attacks in the Problem SpaceFabio Pierazzi, Feargus Pendlebury, Jacopo Cortellazzi, Lorenzo CavallaroS&P 2020 · 334 citations
- Improving Robustness of ML Classifiers against Realizable Evasion Attacks Using Conserved FeaturesLiang Tong, Bo Li, Chen Hajaj, Chaowei Xiao et al.USENIX Security 2019 · 95 citations
- Fast Minimum-norm Adversarial Attacks through Adaptive Norm ConstraintsMaura Pintor, Fabio Roli, Wieland Brendel, Battista BiggioNeurIPS 2021 · 94 citations
Related papers
- Discrete Adversarial Attack to Models of CodeFengjuan Gao, Yu Wang, Ke WangPLDI 2023 · 23 citations
- Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial AttacksMilad Nasr, Yanick Fratantonio, Luca Invernizzi, Ange Albertini et al.CCS 2025
- Adversarial Robustness with Non-uniform PerturbationsEcenaz Erdemir, Jeffrey Bickford, Luca Melis, Sergül AydöreNeurIPS 2021 · 37 citations
- Robust Android Malware Detection against Adversarial Example AttacksHeng Li, Shiyao Zhou, Wei Yuan, Xiapu Luo et al.WWW 2021 · 56 citations
- Adversarial Defense via Learning to Generate Diverse AttacksYunseok Jang, Tianchen Zhao, Seunghoon Hong, Honglak LeeICCV 2019 · 88 citations
