Beyond the Stars: Multimodal Detection of Scams on GitHub
Tillson Galloway, Kevin Valakuzhy, Manos Antonakakis, Fabian Monrose
摘要
The open-source software development ecosystem has become a prime target for phishing and supply-chain attacks, precisely because its openness encourages trust. Defenders fighting these attacks operate in an environment where adversaries generate malicious artifacts at industrial scale, evade signature-based detection through trivial mutations, and exploit delays inherent in community-driven labeling. This combination produces sparse, biased ground truth and extreme class imbalance—conditions under which conventional evaluation pipelines break down. Unsurprisingly, many detection systems that excel offline degrade sharply in real-world deployments, limiting their impact against sustained campaigns. We present OctoWatch , a multimodal machine learning system for expanding detection of large-scale malicious source code repositories under real-world operational constraints while enabling rapid analyst pivoting through learned representations. The system fuses natural-language representations of repository metadata and file contents with structural embeddings of user–repository interactions and temporal activity statistics. Evaluated on real-world datasets—including large-scale secret-exposure attacks, prior threat-intelligence reporting, and a previously unreported phishing campaign—the approach consistently outperforms prior baselines in both precision and Matthews Correlation Coefficient (MCC). To assess operational viability, we conduct an online deployment that processes over ten million newly active repositories and identifies 14,667 total malicious repositories absent from major OSINT feeds. The system further enables discovery of coordinated campaigns, including one activity cluster associated with $5.9 million USD in cryptocurrency transactions on the Ethereum blockchain. Responsible disclosure of these findings led to the takedown of thousands of malicious repositories, demonstrating that multimodal fusion enables scalable detection of software-ecosystem abuse that systematically evades existing community-driven defenses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Reading the Tea leaves: A Comparative Analysis of Threat IntelligenceVector Guo Li, Matthew Dunn, Paul Pearce, Damon McCoy 等USENIX Security 2019 · 被引用 123 次
- Cybercriminal Minds: An investigative study of cryptocurrency abuses in the Dark WebSeunghyeon Lee, Changhoon Yoon, Heedo Kang, Yeonkeun Kim 等NDSS 2019 · 被引用 100 次
- Deep Entity Classification: Abusive Account Detection for Online Social NetworksTeng Xu, Gerard Goossen, Huseyin Kerem Cevahir, Sara Khodeir 等USENIX Security 2021 · 被引用 41 次
- Cybercrime Bitcoin Revenue Estimations: Quantifying the Impact of Methodology and CoverageGibran Gómez, Kevin van Liebergen, Juan CaballeroCCS 2023 · 被引用 10 次
相关 Paper
- Rep2Vec: Repository Embedding via Heterogeneous Graph Adversarial Contrastive LearningYiyue Qian, Yiming Zhang, Qianlong Wen, Yanfang Ye 等KDD 2022 · 被引用 17 次
- CtPhishCapture: Uncovering Credential-Theft-Based Phishing Scams Targeting Cryptocurrency WalletsHui Jiang, Zhenrui Zhang, Xiang Li, Yan Li 等NDSS 2026 · 被引用 1 次
- Six Million (Suspected) Fake Stars on GitHub: A Growing Spiral of Popularity Contests, Spams, and MalwareHao He, Haoqin Yang, Philipp Burckhardt, Alexandros Kapravelos 等ICSE 2026 · 被引用 2 次
- Identifying bad software changes via multimodal anomaly detection for online service systemsNengwen Zhao, Junjie Chen, Zhaoyang Yu, Honglin Wang 等FSE 2021 · 被引用 89 次
- How Bad Can It Git? Characterizing Secret Leakage in Public GitHub RepositoriesMichael Meli, Matthew R. McNiece, Bradley ReavesNDSS 2019 · 被引用 130 次
