Prevalence and Impact of Low-Entropy Packing Schemes in the Malware Ecosystem
Alessandro Mantovani, Simone Aonzo, Xabier Ugarte-Pedrero, Alessio Merlo, Davide Balzarotti
摘要
—An open research problem on malware analysis is how to statically distinguish between packed and non-packed executables. This has an impact on antivirus software and malware analysis systems, which may need to apply different heuristics or to resort to more costly code emulation solutions to deal with the presence of potential packing routines. It can also affect the results of many research studies in which the authors adopt algorithms that are specifically designed for packed or non-packed binaries. Therefore, a wrong answer to the question “is this executable packed?” can make the difference between malware evasion and detection. It has long been known that packing and entropy are strongly correlated, often leading to the wrong assumption that a low entropy score implies that an executable is NOT packed. Exceptions to this rule exist, but they have always been considered as one-off cases, with a negligible impact on any large scale experiment. However, if such an assumption might have been acceptable in the past, our experiments show that this is not the case anymore as an increasing and remarkable number of packed malware samples implement proper schemes to keep their entropy low. In this paper, we empirically investigate and measure this problem by analyzing a dataset of 50K low-entropy Windows malware samples. Our tests show that, despite all samples have a low entropy value, over 30% of them adopt some form of runtime packing. We then extended our analysis beyond the pure entropy, by considering all static features that have been proposed so far to identify packed code. Again, our tests show that even a state of the art machine learning classifier is unable to conclude whether a
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Obfuscation-Resilient Executable Payload Extraction From Packed MalwareBinlin Cheng, Jiang Ming, Erika A. Leal, Haotian Zhang 等USENIX Security 2021 · 被引用 29 次
- SYMBEXCEL: Automated Analysis and Understanding of Malicious Excel 4.0 MacrosNicola Ruaro, Fabio Pagani, Stefano Ortolani, Christopher Kruegel 等S&P 2022 · 被引用 12 次
- PackGenome: Automatically Generating Robust YARA Rules for Accurate Malware Packer DetectionShijia Li, Jiang Ming, Pengda Qiu, Qiyuan Chen 等CCS 2023 · 被引用 10 次
- KEENHash: Hashing Programs into Function-Aware Embeddings for Large-Scale Binary Code Similarity AnalysisZhijie Liu, Qiyi Tang, Sen Nie, Shi Wu 等ISSTA 2025 · 被引用 1 次
- Hiding Critical Program Components via Ambiguous TranslationChijung Jung, Doowon Kim, An Chen, Weihang Wang 等ICSE 2022 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- When Malware is Packin' Heat; Limits of Machine Learning Classifiers Based on Static Analysis FeaturesHojjat Aghakhani, Fabio Gritti, Francesco Mecca, Martina Lindorfer 等NDSS 2020
- Adversarially Robust Assembly Language Model for Packed Executables DetectionShijia Li, Jiang Ming, Lanqing Liu, Longwei Yang 等CCS 2025
- Things You May Not Know About Android (Un)Packers: A Systematic Study based on Whole-System EmulationYue Duan, Mu Zhang, Abhishek Vasisht Bhaskar, Heng Yin 等NDSS 2018 · 被引用 87 次
- Decoding the Secrets of Machine Learning in Malware Classification: A Deep Dive into Datasets, Feature Extraction, and Model PerformanceSavino Dambra, Yufei Han, Simone Aonzo, Platon Kotzias 等CCS 2023 · 被引用 28 次
- Does Every Second Count? Time-based Evolution of Malware Behavior in SandboxesAlexander Küchler, Alessandro Mantovani, Yufei Han, Leyla Bilge 等NDSS 2021
