MalTotal: Cost-Effective and Language-Agnostic Malicious Code Poisoning Detection for Millions of Repositories
Jian Zhao, Shenao Wang, Qingyang Wu, Yanjie Zhao, Xiao Cheng, Haoyu Wang
Abstract
The widespread adoption of open source software (OSS) has introduced significant security risks, with malicious code poisoning attacks increasingly targeting public package registries and open-source platforms. Existing detection approaches, including heuristic-, learning-, and LLM-based methods, suffer from language-specific designs, limited generalization, and high analysis costs, making them unsuitable for large-scale multi-language analysis. To address these challenges, we propose MalTotal, a scalable and cost-effective framework for language-agnostic malicious code detection. MalTotal leverages LLM-assisted semantic reasoning to identify sensitive APIs, perform hybrid semantic slicing, and reconstruct malicious behavior contexts while reducing analysis overhead. Our evaluations show that MalTotal outperforms 8 state-of-the-art baselines, achieving an average F1-score of 93.1% across 5 mainstream languages. Its hybrid slicing reduces LLM token consumption by 94.0%, lowering the analysis cost from 5.19 on 2,168 repositories. In a large-scale study of 120K GitHub repositories containing over 7.3 million files, MalTotal discovered 564 previously unknown malicious repositories across multiple languages at a total cost of $338. These results demonstrate the effectiveness, scalability, and cost-efficiency of MalTotal in mitigating large-scale code poisoning attacks.
CCS Concepts: • Security and privacy → Malware and its mitigation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83aabce8-c11d-4537-acfd-3889b6d0a597Builds on20
- Small World with High Risks: A Study of Security Threats in the npm EcosystemMarkus Zimmermann, Cristian-Alexandru Staicu, Cam Tenny, Michael PradelUSENIX Security 2019 · 281 citations
- Practical Automated Detection of Malicious npm PackagesAdriana Sejfia, Max SchäferICSE 2022 · 65 citations
- DONAPI: Malicious NPM Packages Detector using Behavior Sequence Knowledge MappingCheng Huang, Nannan Wang, Ziyan Wang, Siqi Sun et al.USENIX Security 2024 · 38 citations
- An Empirical Study of Malicious Code In PyPI EcosystemWenbo Guo, Zhengzi Xu, Chengwei Liu, Cheng Huang et al.ASE 2023 · 31 citations
- Malicious Package Detection using Metadata InformationSajal Halder, Michael Bewong, Arash Mahboubi, Yinhao Jiang et al.WWW 2024 · 23 citations
Related papers
- SpiderScan: Practical Detection of Malicious NPM Packages Based on Graph-Based Behavior Modeling and MatchingYiheng Huang, Ruisi Wang, Wen Zheng, Zhuotong Zhou et al.ASE 2024 · 4 citations
- Efficient Code Analysis via Graph Representation Learning-Guided Large Language ModelsHang Gao, Tao Peng, Baoquan Cui, Hong Huang et al.ICML 2026 · 1 citation
- MalGuard: Towards Real-Time, Accurate, and Actionable Detection of Malicious Packages in PyPI EcosystemXingan Gao, Xiaobing Sun, Sicong Cao, Kaifeng Huang et al.USENIX Security 2025
- Maltracker: A Fine-Grained NPM Malware Tracker Copiloted by LLM-Enhanced DatasetZeliang Yu, Ming Wen, Xiaochen Guo, Hai JinISSTA 2024 · 16 citations
- ProfMal: Detecting Malicious NPM Packages by the Synergy between Static and Dynamic AnalysisYiheng Huang, Wen Zheng, Susheng Wu, Bihuan Chen et al.ASE 2025 · 2 citations
