1+1>2: Integrating Deep Code Behaviors with Metadata Features for Malicious PyPI Package Detection
Xiaobing Sun, Xingan Gao, Sicong Cao, Lili Bo, Xiaoxue Wu, Kaifeng Huang
摘要
PyPI, the official package registry for Python, has seen a surge in the number of malicious package uploads in recent years. Prior studies have demonstrated the effectiveness of learning-based solutions in malicious package detection. However, manually-crafted expert rules are expensive and struggle to keep pace with the rapidly evolving malicious behaviors, while deep features automatically extracted from code are still inaccurate in certain cases. To mitigate these issues, in this paper, we propose Ea4mp, a novel approach which integrates deep code behaviors with metadata features to detect malicious PyPI packages. Specifically, Ea4mp extracts code behavior sequences from all script files and fine-tunes a BERT model to learn deep semantic features of malicious code. In addition, we realize the value of metadata information and construct an ensemble classifier to combine the strengths of deep code behavior features and metadata features for more effective detection. We evaluated Ea4mp against three state-of-the-art baselines on a newly constructed dataset. The experimental results show that Ea4mp improves precision by 6.9%-24.6% and recall by 10.5%-18.4%. With Ea4mp, we successfully identified 119 previously unknown malicious packages from a pool of 46,573 newly-uploaded packages over a three-week period, and 82 out of them have been removed by the PyPI official.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- ConfuGuard: Using Metadata to Detect Active and Stealthy Package Confusion Attacks Accurately and at ScaleWenxin Jiang, Berk Çakar, Mikola Lysenko, James C DavisICSE 2026 · 被引用 2 次
- Efficient Code Analysis via Graph Representation Learning-Guided Large Language ModelsHang Gao, Tao Peng, Baoquan Cui, Hong Huang 等ICML 2026 · 被引用 1 次
- Cutting the Gordian Knot: Detecting Malicious PyPI Packages via a Knowledge-Mining FrameworkWenbo Guo, Chengwei Liu, Ming Kang, Yiran Zhang 等USENIX Security 2026 · 被引用 1 次
- MalGuard: Towards Real-Time, Accurate, and Actionable Detection of Malicious Packages in PyPI EcosystemXingan Gao, Xiaobing Sun, Sicong Cao, Kaifeng Huang 等USENIX Security 2025
- MalTotal: Cost-Effective and Language-Agnostic Malicious Code Poisoning Detection for Millions of RepositoriesJian Zhao, Shenao Wang, Qingyang Wu, Yanjie Zhao 等ISSTA 2026
它引用的顶会 Paper11
- MVD: Memory-Related Vulnerability Detection Based on Flow-Sensitive Graph Neural NetworksSicong Cao, Xiaobing Sun, Lili Bo, Rongxin Wu 等ICSE 2022 · 被引用 100 次
- Practical Automated Detection of Malicious npm PackagesAdriana Sejfia, Max SchäferICSE 2022 · 被引用 65 次
- LastPyMile: identifying the discrepancy between sources and packagesDuc-Ly Vu, Fabio Massacci, Ivan Pashchenko, Henrik Plate 等FSE 2021 · 被引用 53 次
- An Empirical Study of Malicious Code In PyPI EcosystemWenbo Guo, Zhengzi Xu, Chengwei Liu, Cheng Huang 等ASE 2023 · 被引用 31 次
- Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection SystemsSicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo 等ICSE 2024 · 被引用 27 次
相关 Paper
- A Needle is an Outlier in a Haystack: Hunting Malicious PyPI Packages with Code ClusteringWentao Liang, Xiang Ling, Jingzheng Wu, Tianyue Luo 等ASE 2023 · 被引用 15 次
- Malicious Package Detection using Metadata InformationSajal Halder, Michael Bewong, Arash Mahboubi, Yinhao Jiang 等WWW 2024 · 被引用 23 次
- Uncovering Similar but Different Packages in PyPI and Potential Security ThreatsSunha Park, Soojin Han, Seunghoon WooFSE 2026
- PyRadar: Towards Automatically Retrieving and Validating Source Code Repository Information for PyPI PackagesKai Gao, Weiwei Xu, Wenhao Yang, Minghui ZhouFSE 2024 · 被引用 7 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
