Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language Model
Side Liu, Jiang Ming, Guodong Zhou, Xinyi Liu, Jianming Fu, Guojun Peng
Abstract
Malicious PDF files have emerged as a persistent threat and become a popular attack vector in web-based attacks. While machine learning-based PDF malware classifiers have shown promise, these classifiers are often susceptible to adversarial attacks, undermining their reliability. To address this issue, recent studies have aimed to enhance the robustness of PDF classifiers. Despite these efforts, the feature engineering underlying these studies remains outdated. Consequently, even with the application of cutting-edge machine learning techniques, these approaches fail to fundamentally resolve the issue of feature instability. To tackle this, we propose a novel approach for PDF feature extraction and PDF malware detection. We introduce the PDFObj IR (PDF Object Intermediate Representation), an assembly-like language framework for PDF objects, from which we extract semantic features using a pretrained language model. Additionally, we construct an Object Reference Graph to capture structural features, drawing inspiration from program analysis. This dual approach enables us to analyze and detect PDF malware based on both semantic and structural features. Experimental results demonstrate that our proposed classifier achieves strong adversarial robustness while maintaining an exceptionally low false positive rate of only 0.07% on baseline dataset compared to state-of-the-art PDF malware classifiers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on15
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler OptimizationSteven H. H. Ding, Benjamin C. M. Fung, Philippe CharlandS&P 2019 · 447 citations
- Transcend: Detecting Concept Drift in Malware Classification ModelsRoberto Jordaney, Kumar Sharad, Santanu Kumar Dash, Zhi Wang et al.USENIX Security 2017 · 325 citations
- Neural Machine Translation Inspired Binary Code Similarity Comparison beyond Function PairsFei Zuo, Xiaopeng Li, Patrick Young, Lannan Luo et al.NDSS 2019 · 262 citations
- Automatically Evading Classifiers: A Case Study on PDF Malware ClassifiersWeilin Xu, Yanjun Qi, David EvansNDSS 2016 · 249 citations
Related papers
- VAPD: An Anomaly Detection Model for PDF Malware Forensics with Adversarial RobustnessSide Liu, Jiang Ming, Yilin Zhou, Jianming Fu et al.USENIX Security 2025
- Improving Robustness of ML Classifiers against Realizable Evasion Attacks Using Conserved FeaturesLiang Tong, Bo Li, Chen Hajaj, Chaowei Xiao et al.USENIX Security 2019 · 95 citations
- On Training Robust PDF Malware ClassifiersYizheng Chen, Shiqi Wang, Dongdong She, Suman JanaUSENIX Security 2020
- MalGraph: Hierarchical Graph Neural Networks for Robust Windows Malware DetectionXiang Ling, Lingfei Wu, Wei Deng, Zhenqing Qu et al.INFOCOM 2022 · 47 citations
- Extract Me If You Can: Abusing PDF Parsers in Malware DetectorsCurtis Carmony, Xunchao Hu, Heng Yin, Abhishek Vasisht Bhaskar et al.NDSS 2016 · 61 citations
