A Needle is an Outlier in a Haystack: Hunting Malicious PyPI Packages with Code Clustering
Wentao Liang, Xiang Ling, Jingzheng Wu, Tianyue Luo, Yanjun Wu
摘要
As the most popular Python software repository, PyPI has become an indispensable part of the Python ecosystem. Regrettably, the open nature of PyPI exposes end-users to substantial security risks stemming from malicious packages. Consequently, the timely and effective identification of malware within the vast number of newly-uploaded PyPI packages has emerged as a pressing concern. Existing detection methods are dependent on difficult-to-obtain explicit knowledge, such as taint sources, sinks, and malicious code patterns, rendering them susceptible to overlooking emergent malicious packages. In this paper, we present a lightweight and effective method, namely MPHunter, to detect malicious packages without requiring any explicit prior knowledge. MPHunter is founded upon two fundamental and insightful observations. First, malicious packages are considerably rarer than benign ones, and second, the functionality of installation scripts for malicious packages diverges significantly from those of benign packages, with the latter frequently forming clusters. Consequently, MPHunter utilizes clustering techniques to group the installation scripts of PyPI packages and identifies outliers. Subsequently, MPHunter ranks the outliers according to their outlierness and the distance between them and known malicious instances, thereby effectively highlighting potential evil packages. With MPHunter, we successfully identified 60 previously unknown malicious packages from a pool of 31,329 newly-uploaded packages over a two-month period. All of them have been confirmed by the PyPI official. Moreover, a manual analysis shows that MPHunter recognizes all potentially malicious installation scripts with a recall of 100% across all analyzed packages. We assert that MPHunter offers a valuable and advantageous supplement to existing detection techniques, augmenting the arsenal of software supply chain security analysis.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper14
- DONAPI: Malicious NPM Packages Detector using Behavior Sequence Knowledge MappingCheng Huang, Nannan Wang, Ziyan Wang, Siqi Sun 等USENIX Security 2024 · 被引用 38 次
- SpiderScan: Practical Detection of Malicious NPM Packages Based on Graph-Based Behavior Modeling and MatchingYiheng Huang, Ruisi Wang, Wen Zheng, Zhuotong Zhou 等ASE 2024 · 被引用 4 次
- 1+1>2: Integrating Deep Code Behaviors with Metadata Features for Malicious PyPI Package DetectionXiaobing Sun, Xingan Gao, Sicong Cao, Lili Bo 等ASE 2024 · 被引用 3 次
- Efficient Code Analysis via Graph Representation Learning-Guided Large Language ModelsHang Gao, Tao Peng, Baoquan Cui, Hong Huang 等ICML 2026 · 被引用 1 次
- Cutting the Gordian Knot: Detecting Malicious PyPI Packages via a Knowledge-Mining FrameworkWenbo Guo, Chengwei Liu, Ming Kang, Yiran Zhang 等USENIX Security 2026 · 被引用 1 次
相关 Paper
- An Empirical Study of Malicious Code In PyPI EcosystemWenbo Guo, Zhengzi Xu, Chengwei Liu, Cheng Huang 等ASE 2023 · 被引用 31 次
- Uncovering Similar but Different Packages in PyPI and Potential Security ThreatsSunha Park, Soojin Han, Seunghoon WooFSE 2026
- MalGuard: Towards Real-Time, Accurate, and Actionable Detection of Malicious Packages in PyPI EcosystemXingan Gao, Xiaobing Sun, Sicong Cao, Kaifeng Huang 等USENIX Security 2025
- PyFEX: Uncovering Evasive Python-based Threats via Resilient and Exhaustive Path ExplorationMeng Wang, Yue Ma, Majid Garoosi, Wenting Fan 等CCS 2026
- Malicious Package Detection using Metadata InformationSajal Halder, Michael Bewong, Arash Mahboubi, Yinhao Jiang 等WWW 2024 · 被引用 23 次
