USENIX Security2020Top-tier venue
On Training Robust PDF Malware Classifiers
Yizheng Chen, Shiqi Wang, Dongdong She, Suman Jana
Abstract
Although state-of-the-art PDF malware classifiers can be trained with almost perfect test accuracy (99%) and extremely low false positive rate (under 0.1%), it has been shown that even a simple adversary can evade them. A practically useful malware classifier must be robust against evasion attacks. However, achieving such robustness is an extremely challenging task. In this paper, we take the first steps towards training robust PDF malware classifiers with verifiable robustness properties. For instance, a robustness property can enforce that no matter how many pages from benign documents are inserted into a PDF malware, the classifier must still classify it as malicious. We demonstrate how the worst-case behavior of a malware classifier with respect to specific robustness properties can be formally verified. Furthermore, we find that training classifiers that satisfy formally verified robustness properties can increase the evasion cost of unbounded (i.e., not bounded by the robustness properties) attackers by eliminating simple evasion attacks. Specifically, we propose a new distance metric that operates on the PDF tree structure and specify two classes of robustness properties including subtree insertions and deletions. We utilize state-of-the-art verifiably robust training method to build robust PDF malware classifiers. Our results show that, we can achieve 92.27% average verified robust accuracy over three properties, while maintaining 99.74% accuracy and 0.56% false positive rate. With simple robustness properties, our robust model maintains 7% higher robust accuracy than all the baseline models against unrestricted whitebox attacks. Moreover, the state-of-the-art and new adaptive evolutionary attackers need up to 10 times larger feature distance and 21 times more PDF basic mutations (e.g., inserting and deleting objects) to evade our robust model than the baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5eb747a3-5fdd-45d7-95f3-08c301ecf312Cited by top-tier papers17
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi et al.USENIX Security 2021 · 241 citations
- Stealing Links from Graph Neural NetworksXinlei He, Jinyuan Jia, Michael Backes, Neil Zhenqiang Gong et al.USENIX Security 2021 · 226 citations
- Robust Android Malware Detection against Adversarial Example AttacksHeng Li, Shiyao Zhou, Wei Yuan, Xiapu Luo et al.WWW 2021 · 56 citations
- RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style TransformationZhen Li, Qian (Guenevere) Chen, Chen Chen, Yayi Zou et al.ICSE 2022 · 39 citations
- Point Cloud Analysis for ML-Based Malicious Traffic Detection: Reducing Majorities of False Positive AlarmsChuanpu Fu, Qi Li, Ke Xu, Jianping WuCCS 2023 · 30 citations
Builds on7
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov et al.S&P 2018 · 987 citations
- Formal Security Analysis of Neural Networks using Symbolic IntervalsShiqi Wang, Kexin Pei, Justin Whitehouse, Junfeng Yang et al.USENIX Security 2018 · 523 citations
- Evading Classifiers by Morphing in the DarkHung Dang, Yue Huang, Ee-Chien ChangCCS 2017 · 125 citations
Related papers
- Improving Robustness of ML Classifiers against Realizable Evasion Attacks Using Conserved FeaturesLiang Tong, Bo Li, Chen Hajaj, Chaowei Xiao et al.USENIX Security 2019 · 95 citations
- Automatically Evading Classifiers: A Case Study on PDF Malware ClassifiersWeilin Xu, Yanjun Qi, David EvansNDSS 2016 · 249 citations
- Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language ModelSide Liu, Jiang Ming, Guodong Zhou, Xinyi Liu et al.CCS 2025
- Adversarial Training for Raw-Binary Malware ClassifiersKeane Lucas, Samruddhi Pai, Weiran Lin, Lujo Bauer et al.USENIX Security 2023
- Learning Security Classifiers with Verified Global Robustness PropertiesYizheng Chen, Shiqi Wang, Yue Qin, Xiaojing Liao et al.CCS 2021 · 26 citations
