TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time
Feargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder, Lorenzo Cavallaro
摘要
Is Android malware classification a solved problem? Published F1 scores of up to 0.99 appear to leave very little room for improvement. In this paper, we argue that results are commonly inflated due to two pervasive sources of experimental bias: "spatial bias" caused by distributions of training and testing data that are not representative of a real-world deployment; and "temporal bias" caused by incorrect time splits of training and testing sets, leading to impossible configurations. We propose a set of space and time constraints for experiment design that eliminates both sources of bias. We introduce a new metric that summarizes the expected robustness of a classifier in a real-world setting, and we present an algorithm to tune its performance. Finally, we demonstrate how this allows us to evaluate mitigation strategies for time decay such as active learning. We have implemented our solutions in TESSERACT, an open source evaluation framework for comparing malware classifiers in a realistic setting. We used TESSERACT to evaluate three Android malware classifiers from the literature on a dataset of 129K applications spanning over three years. Our evaluation confirms that earlier published results are biased, while also revealing counter-intuitive performance and showing that appropriate tuning can lead to significant improvements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper62
- Intriguing Properties of Adversarial ML Attacks in the Problem SpaceFabio Pierazzi, Feargus Pendlebury, Jacopo Cortellazzi, Lorenzo CavallaroS&P 2020 · 被引用 334 次
- CADE: Detecting and Explaining Concept Drift Samples for Security ApplicationsLimin Yang, Wenbo Guo, Qingying Hao, Arridhana Ciptadi 等USENIX Security 2021 · 被引用 241 次
- Enhancing State-of-the-art Classifiers with API Semantics to Detect Evolved Android MalwareXiaohan Zhang, Yuan Zhang, Ming Zhong, Daizong Ding 等CCS 2020 · 被引用 173 次
- Transcending TRANSCEND: Revisiting Malware Classification in the Presence of Concept DriftFederico Barbero, Feargus Pendlebury, Fabio Pierazzi, Lorenzo CavallaroS&P 2022 · 被引用 124 次
- Detecting and Characterizing Lateral Phishing at ScaleGrant Ho, Asaf Cidon, Lior Gavish, Marco Schweighauser 等USENIX Security 2019 · 被引用 113 次
它引用的顶会 Paper3
- MaMaDroid: Detecting Android Malware by Building Markov Chains of Behavioral ModelsEnrico Mariconti, Lucky Onwuzurike, Panagiotis Andriotis, Emiliano De Cristofaro 等NDSS 2017 · 被引用 471 次
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su 等CCS 2018 · 被引用 336 次
- Transcend: Detecting Concept Drift in Malware Classification ModelsRoberto Jordaney, Kumar Sharad, Santanu Kumar Dash, Zhi Wang 等USENIX Security 2017 · 被引用 325 次
相关 Paper
- Continuous Learning for Android Malware DetectionYizheng Chen, Zhoujie Ding, David A. WagnerUSENIX Security 2023
- The Illusion of Success: Learning-Based Android Malware Detectors (Replicability Study)Michael Tegegn, Julia RubinISSTA 2026
- DRMD: Deep Reinforcement Learning for Malware Detection Under Concept DriftShae McFadden, Myles Foley, Mario D'Onghia, Chris Hicks 等AAAI 2026 · 被引用 7 次
- Spotless Sandboxes: Evading Malware Analysis Systems Using Wear-and-Tear ArtifactsNajmeh Miramirkhani, Mahathi Priya Appini, Nick Nikiforakis, Michalis PolychronakisS&P 2017 · 被引用 134 次
- Harvesting Runtime Values in Android Applications That Feature Anti-Analysis TechniquesSiegfried Rasthofer, Steven Arzt, Marc Miltenberger, Eric BoddenNDSS 2016 · 被引用 157 次
