ALMOND: Learning an Assembly Language Model for 0-Shot Code Obfuscation Detection
Xuezixiang Li, Sheng Yu, Heng Yin
Abstract
Code obfuscation is a technique used to protect software by making it difficult to understand and reverse engineer. However, it can also be exploited for malicious purposes such as code plagiarism or developing malicious programs. Learning-based techniques have achieved great success with the help of supervised learning and labeled training sets. However, when faced with real-life environments involving privately developed and undisclosed obfuscators, these supervised learning methods often raise concerns about generalizability and robustness when facing unseen and unknown classes of obfuscation techniques. This paper presents ALMOND, a novel zero-shot approach for detecting code obfuscation in binary executables. Unlike previous supervised learning methods, ALMOND does not require labeled obfuscated samples for training. Instead, it leverages a language model pre-trained only on unobfuscated assembly code to identify the linguistic deviations introduced by obfuscation. The key innovation is the use of "error-perplexity" as a detection metric, which focuses on tokens the model fails to predict. Continuous Error Perplexity further enhances this to capture consecutive prediction errors characteristic of obfuscated sequences. Experiments show ALMOND achieves 96.3% accuracy on unseen obfuscation methods, outperforming supervised baselines. On real-world malware samples, it demonstrates an AUC of 0.869 and significantly outperforms the supervise-learning baseline. Our Dataset, pre-trained model, and code of evaluation will be available at https://github.com/palmtreemodel/ALMOND CCS Concepts: • Security and privacy → Software reverse engineering; Intrusion/anomaly detection and malware mitigation; • Theory of computation → Program analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ef6f78f-7292-4fee-b796-8c930d67d70bBuilds on11
- DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep LearningMin Du, Feifei Li, Guineng Zheng, Vivek SrikumarCCS 2017 · 1,823 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler OptimizationSteven H. H. Ding, Benjamin C. M. Fung, Philippe CharlandS&P 2019 · 447 citations
- Order Matters: Semantic-Aware Neural Networks for Binary Code Similarity DetectionZeping Yu, Rui Cao, Qiyi Tang, Sen Nie et al.AAAI 2020 · 265 citations
- Log-based Anomaly Detection with Deep Learning: How Far Are We?Van-Hoang Le, Hongyu ZhangICSE 2022 · 212 citations
Related papers
- Adversarially Robust Assembly Language Model for Packed Executables DetectionShijia Li, Jiang Ming, Lanqing Liu, Longwei Yang et al.CCS 2025
- CLAP: Learning Transferable Binary Code Representations with Natural Language SupervisionHao Wang, Zeyu Gao, Chao Zhang, Zihan Sha et al.ISSTA 2024 · 32 citations
- Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code ObfuscationSeyedreza Mohseni, Seyedali Mohammadi, Deepa Tilwani, Yash Saxena et al.AAAI 2025 · 6 citations
- DOBF: A Deobfuscation Pre-Training Objective for Programming LanguagesMarie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume LampleNeurIPS 2021 · 174 citations
- Tackling runtime-based obfuscation in Android with TIROMichelle Y. Wong, David LieUSENIX Security 2018 · 59 citations
