ALMOND: Learning an Assembly Language Model for 0-Shot Code Obfuscation Detection
Xuezixiang Li, Sheng Yu, Heng Yin
摘要
Code obfuscation is a technique used to protect software by making it difficult to understand and reverse engineer. However, it can also be exploited for malicious purposes such as code plagiarism or developing malicious programs. Learning-based techniques have achieved great success with the help of supervised learning and labeled training sets. However, when faced with real-life environments involving privately developed and undisclosed obfuscators, these supervised learning methods often raise concerns about generalizability and robustness when facing unseen and unknown classes of obfuscation techniques. This paper presents ALMOND, a novel zero-shot approach for detecting code obfuscation in binary executables. Unlike previous supervised learning methods, ALMOND does not require labeled obfuscated samples for training. Instead, it leverages a language model pre-trained only on unobfuscated assembly code to identify the linguistic deviations introduced by obfuscation. The key innovation is the use of "error-perplexity" as a detection metric, which focuses on tokens the model fails to predict. Continuous Error Perplexity further enhances this to capture consecutive prediction errors characteristic of obfuscated sequences. Experiments show ALMOND achieves 96.3% accuracy on unseen obfuscation methods, outperforming supervised baselines. On real-world malware samples, it demonstrates an AUC of 0.869 and significantly outperforms the supervise-learning baseline. Our Dataset, pre-trained model, and code of evaluation will be available at https://github.com/palmtreemodel/ALMOND CCS Concepts: • Security and privacy → Software reverse engineering; Intrusion/anomaly detection and malware mitigation; • Theory of computation → Program analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 被引用 1,823 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler OptimizationSteven H. H. Ding, Benjamin C. M. Fung, Philippe CharlandS&P 2019 · 被引用 447 次
- Order Matters: Semantic-Aware Neural Networks for Binary Code Similarity DetectionZeping Yu, Rui Cao, Qiyi Tang, Sen Nie 等AAAI 2020 · 被引用 265 次
- Log-based Anomaly Detection with Deep Learning: How Far Are We?Van-Hoang Le, Hongyu ZhangICSE 2022 · 被引用 212 次
相关 Paper
- Adversarially Robust Assembly Language Model for Packed Executables DetectionShijia Li, Jiang Ming, Lanqing Liu, Longwei Yang 等CCS 2025
- CLAP: Learning Transferable Binary Code Representations with Natural Language SupervisionHao Wang, Zeyu Gao, Chao Zhang, Zihan Sha 等ISSTA 2024 · 被引用 32 次
- Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code ObfuscationSeyedreza Mohseni, Seyedali Mohammadi, Deepa Tilwani, Yash Saxena 等AAAI 2025 · 被引用 6 次
- DOBF: A Deobfuscation Pre-Training Objective for Programming LanguagesMarie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume LampleNeurIPS 2021 · 被引用 174 次
- Tackling runtime-based obfuscation in Android with TIROMichelle Y. Wong, David LieUSENIX Security 2018 · 被引用 59 次
