USENIX Security2023Top-tier venue
Beyond Typosquatting: An In-depth Look at Package Confusion
Shradha Neupane, Grant Holmes, Elizabeth Wyss, Drew Davidson, Lorenzo De Carli
Abstract
Package confusion incidents-where a developer is misled into importing a package other than the intended one-are one of the most severe issues in supply chain security with significant security implications, especially when the wrong package has malicious functionality. While the prevalence of the issue is generally well-documented, little work has studied the range of mechanisms by which confusion in a package name could arise or be employed by an adversary. In our work, we present the first comprehensive categorization of the mechanisms used to induce confusion, and we show how this understanding can be used for detection. First, we use qualitative analysis to identify and rigorously define 13 categories of confusion mechanisms based on a dataset of 1200+ documented attacks. Results show that, while package confusion is thought to mostly exploit typing errors, in practice attackers use a variety of mechanisms, many of which work at semantic, rather than syntactic, level. Equipped with our categorization, we then define detectors for the discovered attack categories, and we evaluate them on the entire npm package set. Evaluation of a sample, performed through an online survey, identifies a subset of highly effective detection rules which (i) return high-quality matches (77% matches marked as potentially or highly confusing, and 18% highly confusing) and (ii) generate low warning overhead (1 warning per 100M+ package pairs). Comparison with state-of-the-art reveals that the large majority of such pairs are not flagged by existing tools. Thus, our work has the potential to concretely improve the identification of confusable package names in the wild.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- DONAPI: Malicious NPM Packages Detector using Behavior Sequence Knowledge MappingCheng Huang, Nannan Wang, Ziyan Wang, Siqi Sun et al.USENIX Security 2024 · 38 citations
- The Imitation Game: Exploring Brand Impersonation Attacks on Social Media PlatformsBhupendra Acharya, Dario Lazzaro, Efrén López-Morales, Adam Oest et al.USENIX Security 2024 · 7 citations
- 1+1>2: Integrating Deep Code Behaviors with Metadata Features for Malicious PyPI Package DetectionXiaobing Sun, Xingan Gao, Sicong Cao, Lili Bo et al.ASE 2024 · 3 citations
- A First Look at Security and Privacy Risks in the RapidAPI EcosystemSong Liao, Long Cheng, Xiapu Luo, Zheng Song et al.CCS 2024 · 3 citations
- ConfuGuard: Using Metadata to Detect Active and Stealthy Package Confusion Attacks Accurately and at ScaleWenxin Jiang, Berk Çakar, Mikola Lysenko, James C DavisICSE 2026 · 2 citations
Builds on8
- Small World with High Risks: A Study of Security Threats in the npm EcosystemMarkus Zimmermann, Cristian-Alexandru Staicu, Cam Tenny, Michael PradelUSENIX Security 2019 · 281 citations
- CHAINIAC: Proactive Software-Update Transparency via Collectively Signed Skipchains and Verified BuildsKirill Nikitin, Eleftherios Kokoris-Kogias, Philipp Jovanovic, Nicolas Gailly et al.USENIX Security 2017 · 144 citations
- Practical Automated Detection of Malicious npm PackagesAdriana Sejfia, Max SchäferICSE 2022 · 65 citations
- LastPyMile: identifying the discrepancy between sources and packagesDuc-Ly Vu, Fabio Massacci, Ivan Pashchenko, Henrik Plate et al.FSE 2021 · 53 citations
- Mobile App SquattingYangyu Hu, Haoyu Wang, Ren He, Li Li et al.WWW 2020 · 44 citations
Related papers
- Towards Measuring Supply Chain Attacks on Package Managers for Interpreted LanguagesRuian Duan, Omar Alrawi, Ranjita Pai Kasturi, Ryan Elder et al.NDSS 2021
- ProfMal: Detecting Malicious NPM Packages by the Synergy between Static and Dynamic AnalysisYiheng Huang, Wen Zheng, Susheng Wu, Bihuan Chen et al.ASE 2025 · 2 citations
- Maltracker: A Fine-Grained NPM Malware Tracker Copiloted by LLM-Enhanced DatasetZeliang Yu, Ming Wen, Xiaochen Guo, Hai JinISSTA 2024 · 16 citations
- What the Fork? Finding Hidden Code Clones in npmElizabeth Wyss, Lorenzo De Carli, Drew DavidsonICSE 2022 · 10 citations
- Jack-in-the-box: An Empirical Study of JavaScript Bundling on the Web and its Security ImplicationsJeremy Rack, Cristian-Alexandru StaicuCCS 2023 · 11 citations
