Lune

USENIX Security2023Top-tier venue

Adversarial Training for Raw-Binary Malware Classifiers

Keane Lucas, Samruddhi Pai, Weiran Lin, Lujo Bauer, Michael K. Reiter, Mahmood Sharif

2023Year
7Top-tier citations

Abstract

Machine learning (ML) models have shown promise in classifying raw executable files (binaries) as malicious or benign with high accuracy. This has led to the increasing influence of ML-based classification methods in academic and real-world malware detection, a critical tool in cybersecurity. However, previous work provoked caution by creating variants of malicious binaries, referred to as adversarial examples, that are transformed in a functionality-preserving way to evade detection. In this work, we investigate the effectiveness of using adversarial training methods to create malware-classification models that are more robust to some state-of-the-art attacks. To train our most robust models, we significantly increase the efficiency and scale of creating adversarial examples to make adversarial training practical, which has not been done before in raw-binary malware detectors. We then analyze the effects of varying the length of adversarial training, as well as analyze the effects of training with various types of attacks. We find that data augmentation does not deter state-of-the-art attacks, but that using a generic gradient-guided method, used in other discrete domains, does improve robustness. We also show that in most cases, models can be made more robust to malware-domain attacks by adversarially training them with lower-effort versions of the same attack. In the best case, we reduce one state-of-the-art attack's success rate from 90% to 5%. We also find that training with some types of attacks can increase robustness to other types of attacks. Finally, we discuss insights gained from our results, and how they can be used to more effectively train robust malware detectors. cluding vastly increasing the number of binaries eligible to be turned into adversarial examples (200 -→ 126,009); code-level optimizations to the adversarial-example-generation code to speed up each individual attack (Sec. 3.3); a distributed system of over 140 workers across thirteen servers to produce adversarial examples in parallel (Sec. 3.1); and training with less effective but computationally cheaper versions (Sec. 3.3) of the originally proposed attacks [33, 36] . Our findings include the following: • Adversarial training using data-augmentation techniques is ineffective in defending against adversarial examples in the domain of raw-binary classification. • In contrast, adversarial training using state-of-the-art attacks can yield robust models (90% -→ 5%, 26% -→ 6%, and 84% -→ 30% attack success rate for Disp-, IPR-, Kreuk-based attacks, respectively) when evaluated against the same type of attack used during training. We also find that these models can reduce success for attacks they were not trained on by up to 65%. We investigate how changing parameters of the attacks used in training (e.g., number of attack iterations, permitted file size increase, type of attack) affects the robustness of the resulting model. • We construct, train with, and evaluate a perturbation strategy inspired by other discrete domains [61, 62] , referred to as Greedy, which modifies the discrete bytes of the input binary without regard to the underlying structure, to insert evasive bytes in the most important locations. We find that this approach is effective in inducing some robustness against all attacks. • Finally, we adversarially train a classifier with adversarial examples created by three different attacks, and find that this classifier exhibits close to best-case robustness against every attack type. Background and Related Work Our work builds on research using deep neural networks (DNNs) for malware detection from raw binaries [32, 46] , along with work generating adversarial examples in the malware domain [18, 19, 33, 36, 52 ] and other domains [5, 7, 9, 23, 41, 55] . Next, we cover the background of these topics, and describe the state of the art in creating adversarial examples for malware detectors and defending against them. DNNs for Malware Detection Various feature types have been proposed for malware detection. Expert-designed features (e.g., [3, 4, 20, 22, 26, 30, 49] ) include vectors describing imported libraries, library or API function calls, network addresses contacted, number of system calls, byte entropy, whether specific strings are present, etc. Hand-crafting these features is time-consuming, but their use in DNNs has comparable results to DNNs that learn directly from raw bytes [3, 32, 36, 46] . In this paper, we only consider DNNs that infer maliciousness directly from raw bytes. MalConv [46] and AvastNet [32] are examples of DNN architectures that use raw-bytes as input. These architectures, used in many related works [18, 21, 36, 52] , achieved 98.5% and 98.6% test accuracies respectively in discriminating unseen executables [36] . Real-world malware detection consists of an ensemble of detection mechanisms, not just including the static analysis these DNNs perform [56] . Other analyses include dynamic analyses of

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers7

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines