Extract Me If You Can: Abusing PDF Parsers in Malware Detectors
Curtis Carmony, Xunchao Hu, Heng Yin, Abhishek Vasisht Bhaskar, Mu Zhang
摘要
Owing to the popularity of the PDF format and the continued exploitation of Adobe Reader, the detection of malicious PDFs remains a concern. All existing detection techniques rely on the PDF parser to a certain extent, while the complexity of the PDF format leaves an abundant space for parser confusion. To quantify the difference between these parsers and Adobe Reader, we create a reference JavaScript extractor by directly tapping into Adobe Reader at locations identified through a mostly automatic binary analysis technique. By comparing the output of this reference extractor against that of several opensource JavaScript extractors on a large data set obtained from VirusTotal, we are able to identify hundreds of samples which existing extractors fail to extract JavaScript from. By analyzing these samples we are able to identify several weaknesses in each of these extractors. Based on these lessons, we apply several obfuscations on a malicious PDF sample, which can successfully evade all the malware detectors tested. We call this evasion technique a PDF parser confusion attack. Lastly, we demonstrate that the reference JavaScript extractor improves the accuracy of existing JavaScript-based classifiers and how it can be used to mitigate these parser limitations in a real-world setting. Permission to freely reproduce all or part of this paper for noncommercial purposes is granted provided that copies bear this notice and the full citation on the first page. Reproduction for commercial purposes is strictly prohibited without the prior written consent of the Internet Society, the first-named author (for reproduction of an entire paper only), and the author's employer if the paper was prepared within the scope of employment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- 1 Trillion Dollar Refund: How To Spoof PDF SignaturesVladislav Mladenov, Christian Mainka, Karsten Meyer zu Selhausen, Martin Grothe 等CCS 2019 · 被引用 23 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- Practical Decryption exFiltration: Breaking PDF EncryptionJens Müller, Fabian Ising, Vladislav Mladenov, Christian Mainka 等CCS 2019 · 被引用 18 次
- Themis: Ambiguity-Aware Network Intrusion Detection based on Symbolic Model ComparisonZhongjie Wang, Shitong Zhu, Keyu Man, Pengxiong Zhu 等CCS 2021 · 被引用 6 次
- Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment DetectorsJiahe Zhang, Jianjun Chen, Qi Wang, Hangyu Zhang 等CCS 2024 · 被引用 2 次
相关 Paper
- Wobfuscator: Obfuscating JavaScript Malware via Opportunistic Translation to WebAssemblyAlan Romano, Daniel Lehmann, Michael Pradel, Weihang WangS&P 2022 · 被引用 40 次
- VAPD: An Anomaly Detection Model for PDF Malware Forensics with Adversarial RobustnessSide Liu, Jiang Ming, Yilin Zhou, Jianming Fu 等USENIX Security 2025
- Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language ModelSide Liu, Jiang Ming, Guodong Zhou, Xinyi Liu 等CCS 2025
- An Empirical Study on the Effects of Obfuscation on Static Machine Learning-Based Malicious JavaScript DetectorsKunlun Ren, Weizhong Qiang, Yueming Wu, Yi Zhou 等ISSTA 2023 · 被引用 11 次
- FV8: A Forced Execution JavaScript Engine for Detecting Evasive TechniquesNikolaos Pantelaios, Alexandros KapravelosUSENIX Security 2024 · 被引用 6 次
