Does Every Second Count? Time-based Evolution of Malware Behavior in Sandboxes
Alexander Küchler, Alessandro Mantovani, Yufei Han, Leyla Bilge, Davide Balzarotti
摘要
—The amount of time in which a sample is executed is one of the key parameters of a malware analysis sandbox. Setting the threshold too high hinders the scalability and reduces the number of samples that can be analyzed in a day; too low and the samples may not have the time to show their malicious behavior, thus reducing the amount and quality of the collected data. Therefore, an analyst needs to find the ‘sweet spot’ that allows to collect only the minimum amount of information required to properly classify each sample. Anything more is wasting resources, anything less is jeopardizing the experiments. Despite its importance, there are no clear guidelines on how to choose this parameter, nor experiments that can help companies to assess the pros and cons of a choice over another. To fill this gap, in this paper we provide the first large-scale study of the impact that the execution time has on both the amount and the quality of the collected events. We measure the evolution of system calls and code coverage, to draw a precise picture of the fraction of runtime behavior we can expect to observe in a sandbox. Finally, we implemented a machine learning based malware detection method, and applied it to the data collected in different time windows, to also report on the relevance of the events observed at different points in time. Our results show that most samples run for either less than two minutes or for more than ten. However, most of the behavior (and 98% of the executed basic blocks) are observed during the first two minutes of execution, which is also the time windows that result in a higher accuracy of our ML classifier. We believe this information can help future researchers and industrial sandboxes to better tune their analysis systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- API2Vec: Learning Representations of API Sequences for Malware DetectionLei Cui, Jiancong Cui, Yuede Ji, Zhiyu Hao 等ISSTA 2023 · 被引用 37 次
- Decoding the Secrets of Machine Learning in Malware Classification: A Deep Dive into Datasets, Feature Extraction, and Model PerformanceSavino Dambra, Yufei Han, Simone Aonzo, Platon Kotzias 等CCS 2023 · 被引用 28 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- SYMBEXCEL: Automated Analysis and Understanding of Malicious Excel 4.0 MacrosNicola Ruaro, Fabio Pagani, Stefano Ortolani, Christopher Kruegel 等S&P 2022 · 被引用 12 次
- Unveiling BYOVD Threats: Malware's Use and Abuse of Kernel DriversAndrea Monzani, Antonio Parata, Andrea Oliveri, Simone Aonzo 等NDSS 2026 · 被引用 5 次
它引用的顶会 Paper5
- Understanding Linux MalwareEmanuele Cozzi, Mariano Graziano, Yanick Fratantonio, Davide BalzarottiS&P 2018 · 被引用 203 次
- Spotless Sandboxes: Evading Malware Analysis Systems Using Wear-and-Tear ArtifactsNajmeh Miramirkhani, Mahathi Priya Appini, Nick Nikiforakis, Michalis PolychronakisS&P 2017 · 被引用 134 次
- Capturing Malware Propagations with Code Injections and Code-Reuse AttacksDavid Korczynski, Heng YinCCS 2017 · 被引用 57 次
- Methodologies for Quantifying (Re-)randomization Security and Timing under JIT-ROPSalman Ahmed, Ya Xiao, Kevin Z. Snow, Gang Tan 等CCS 2020 · 被引用 21 次
- Prevalence and Impact of Low-Entropy Packing Schemes in the Malware EcosystemAlessandro Mantovani, Simone Aonzo, Xabier Ugarte-Pedrero, Alessio Merlo 等NDSS 2020
相关 Paper
- A Lustrum of Malware Network Communication: Evolution and InsightsChaz Lever, Platon Kotzias, Davide Balzarotti, Juan Caballero 等S&P 2017 · 被引用 86 次
- When Malware Changed Its Mind: An Empirical Study of Variable Program Behaviors in the Real WorldErin Avllazagaj, Ziyun Zhu, Leyla Bilge, Davide Balzarotti 等USENIX Security 2021 · 被引用 1 次
- Humans vs. Machines in Malware ClassificationSimone Aonzo, Yufei Han, Alessandro Mantovani, Davide BalzarottiUSENIX Security 2023
- When Malware is Packin' Heat; Limits of Machine Learning Classifiers Based on Static Analysis FeaturesHojjat Aghakhani, Fabio Gritti, Francesco Mecca, Martina Lindorfer 等NDSS 2020
- Training Robust ML-based Raw-Binary Malware Detectors in Hours, not MonthsKeane Lucas, Weiran Lin, Lujo Bauer, Michael K. Reiter 等CCS 2024 · 被引用 2 次
