Maltracker: A Fine-Grained NPM Malware Tracker Copiloted by LLM-Enhanced Dataset
Zeliang Yu, Ming Wen, Xiaochen Guo, Hai Jin
Abstract
As the largest package registry, Node Package Manager (NPM) has become the prime target for various supply chain attacks recently and has been flooded with numerous malicious packages, posing significant security risks to end-users. Learning-based methods have demonstrated promising performance with good adaptability to various types of attacks. However, they suffer from two main limitations. First, they often utilize metadata features or coarse-grained code features extracted at the package level while overlooking complex code semantics. Second, the dataset used to train the model often suffers from a lack of variety both in quantity and diversity, and thus cannot detect significant types of attacks. To address these problems, we introduce Maltracker, a learningbased NPM malware tracker based on fine-grained features empowered by LLM-enhanced dataset. First, Maltracker constructs precise call graphs to extract suspicious functions that are reachable to a pre-defined set of sensitive APIs, and then utilizes community detection algorithm to identify suspicious code gadgets based on program dependency graph, from which fine-grained features are then extracted. To address the second limitation, we extend the dataset using advanced large language models (LLM) to translate malicious functions from other languages (e.g., C/C++, Python, and Go) into JavaScript. Evaluations shows that Maltracker can achieve an improvement of about 12.6% in terms of F1-score at the package level and 31.0% at the function level compared with the SOTA learning-based methods. Moreover, the key components of 𝑀𝑎𝑙𝑡𝑟𝑎𝑐𝑘𝑒𝑟 all contribute to the effectiveness of its performance. Finally, Maltracker has also detected 230 new malicious packages in NPM and received 61 thanks letters, among which some contain new malicious behaviors that cannot be detected by existing tools.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5f3ace99-9ca5-4748-b835-ca45152745e8Cited by top-tier papers8
- Efficient Code Analysis via Graph Representation Learning-Guided Large Language ModelsHang Gao, Tao Peng, Baoquan Cui, Hong Huang et al.ICML 2026 · 1 citation
- MalGuard: Towards Real-Time, Accurate, and Actionable Detection of Malicious Packages in PyPI EcosystemXingan Gao, Xiaobing Sun, Sicong Cao, Kaifeng Huang et al.USENIX Security 2025
- "I wasn't sure if this is indeed a security risk": Data-driven Understanding of Security Issue Reporting in GitHub Repositories of Open Source npm PackagesRajdeep Ghosh, Shiladitya De, Mainack MondalUSENIX Security 2025
- Fact-Aligned and Template-Constrained Static Analyzer Rule Enhancement with LLMsZongze Jiang, Ming Wen, Ge Wen, Hai JinASE 2025
- Bridging Expert Reasoning and LLM Detection: A Knowledge-Driven Framework for Malicious PackagesWenbo Guo, Shiwen Song, Jiaxun Guo, Zhengzi Xu et al.WWW 2026
Related papers
- ProfMal: Detecting Malicious NPM Packages by the Synergy between Static and Dynamic AnalysisYiheng Huang, Wen Zheng, Susheng Wu, Bihuan Chen et al.ASE 2025 · 2 citations
- SpiderScan: Practical Detection of Malicious NPM Packages Based on Graph-Based Behavior Modeling and MatchingYiheng Huang, Ruisi Wang, Wen Zheng, Zhuotong Zhou et al.ASE 2024 · 4 citations
- From Noise to Signal: Precisely Identify Affected Packages of Known Vulnerabilities in npm EcosystemYingyuan Pu, Lingyun Ying, Yacong GuNDSS 2026 · 4 citations
- Towards Measuring Supply Chain Attacks on Package Managers for Interpreted LanguagesRuian Duan, Omar Alrawi, Ranjita Pai Kasturi, Ryan Elder et al.NDSS 2021
- Cutting the Gordian Knot: Detecting Malicious PyPI Packages via a Knowledge-Mining FrameworkWenbo Guo, Chengwei Liu, Ming Kang, Yiran Zhang et al.USENIX Security 2026 · 1 citation
