FeatureSmith: Automatically Engineering Features for Malware Detection by Mining the Security Literature
Ziyun Zhu, Tudor Dumitras
摘要
Malware detection increasingly relies on machine learning techniques, which utilize multiple features to separate the malware from the benign apps. The e↵ectiveness of these techniques primarily depends on the manual feature engineering process, based on human knowledge and intuition. However, given the adversaries' e↵orts to evade detection and the growing volume of publications on malware behaviors, the feature engineering process likely draws from a fraction of the relevant knowledge. We propose an end-to-end approach for automatic feature engineering. We describe techniques for mining documents written in natural language (e.g. scientific papers) and for representing and querying the knowledge about malware in a way that mirrors the human feature engineering process. Specifically, we first identify abstract behaviors that are associated with malware, and then we map these behaviors to concrete features that can be tested experimentally. We implement these ideas in a system called FeatureSmith, which generates a feature set for detecting Android malware. We train a classifier using these features on a large data set of benign and malicious apps. This classifier achieves a 92.5% true positive rate with only 1% false positives, which is comparable to the performance of a state-of-the-art Android malware detector that relies on manually engineered features. In addition, FeatureSmith is able to suggest informative features that are absent from the manually engineered set and to link the features generated to abstract concepts that describe malware behaviors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Towards the Detection of Inconsistencies in Public Security Vulnerability ReportsYing Dong, Wenbo Guo, Yueqi Chen, Xinyu Xing 等USENIX Security 2019 · 被引用 149 次
- Understanding the Reproducibility of Crowd-reported Security VulnerabilitiesDongliang Mu, Alejandro Cuevas, Limin Yang, Hang Hu 等USENIX Security 2018 · 被引用 138 次
- Understanding and Securing Device Vulnerabilities through Automated Bug Report AnalysisXuan Feng, Xiaojing Liao, XiaoFeng Wang, Haining Wang 等USENIX Security 2019 · 被引用 46 次
- YARIX: Scalable YARA-based Malware IntelligenceMichael Brengel, Christian RossowUSENIX Security 2021 · 被引用 23 次
- Demystifying the Vetting Process of Voice-controlled Skills on MarketsDawei Wang, Kai Chen, Wei WangUbiComp 2021 · 被引用 12 次
它引用的顶会 Paper1
相关 Paper
- The Illusion of Success: Learning-Based Android Malware Detectors (Replicability Study)Michael Tegegn, Julia RubinISSTA 2026
- Enhancing State-of-the-art Classifiers with API Semantics to Detect Evolved Android MalwareXiaohan Zhang, Yuan Zhang, Ming Zhong, Daizong Ding 等CCS 2020 · 被引用 173 次
- Continuous Learning for Android Malware DetectionYizheng Chen, Zhoujie Ding, David A. WagnerUSENIX Security 2023
- Guided Retraining to Enhance the Detection of Difficult Android MalwareNadia Daoudi, Kevin Allix, Tegawendé F. Bissyandé, Jacques KleinISSTA 2023 · 被引用 4 次
- MaMaDroid: Detecting Android Malware by Building Markov Chains of Behavioral ModelsEnrico Mariconti, Lucky Onwuzurike, Panagiotis Andriotis, Emiliano De Cristofaro 等NDSS 2017 · 被引用 471 次
