SIRAJ: A Unified Framework for Aggregation of Malicious Entity Detectors
Saravanan Thirumuruganathan, Mohamed Nabeel, Euijin Choo, Issa Khalil, Ting Yu
Abstract
High-quality intelligence of Internet threat (e.g., malware files, malicious domains, phishing URLs and malicious IPs) are important for both security practitioners and the research community. Given the agility of attackers, the scale of the Internet, and the fast-evolving landscape of threats, one could not rely solely on a single source (such as an anti-malware engine or an IP blacklist) for obtaining accurate, up-to-date, and comprehensive threat analysis. Instead, we need to aggregate the analysis from multiple sources. However, it is non-trivial to do such aggregation effectively. A common practice is to label an indicator (malware, domains, URLs, etc.) as malicious if it is marked by a number of sources above an ad-hoc certain threshold. Often, this results in sub-optimal performance as it assumes that all sources are of similar quality/expertise, independent, and temporally stable, which unfortunately are often not true in practice. A natural alternative is to train a supervised machine learning model. However, this approach needs a sufficiently large amount of manually labeled ground truth, which is time-consuming to collect and has to be updated frequently, resulting in substantial recurring costs. In this paper, we propose SIRAJ, a novel framework for aggregating the detection output of various intelligence sources such as anti-malware engines. SIRAJ is based on the pretrain and fine-tune paradigm. Specifically, we use self-supervised learning-based approaches to learn a pre-trained embedding model that converts multi-source inputs into a high-dimensional embedding. The embeddings are learned through three carefully designed pretext tasks that imbue them with knowledge about dependencies between scanners and their temporal dynamics. The learned embeddings could be used for diverse downstream machine learning tasks. SIRAJ is designed to be general and can be used for diverse domains such as URLs, malware, and IPs. Further, SIRAJ works well even when there is limited to no labeled data available. Through extensive experiments, we show that our learned representations can produce results comparable to supervised methods while only requiring as little as 100 labeled samples. Importantly, the results show that SIRAJ accurately detects threat indicators much earlier than the baseline algorithms, a feat that is critical against short-lived indicators like Phishing URLs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9f39fe7-3d14-4181-89b4-d2f67d65c435Cited by top-tier papers6
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Detecting and Mitigating Sampling Bias in Cybersecurity with Unlabeled DataSaravanan Thirumuruganathan, Fatih Deniz, Issa Khalil, Ting Yu et al.USENIX Security 2024 · 6 citations
- Models on the Move: Towards Feasible Embedded AI for Intrusion Detection on Vehicular CAN BusHe Xu, Di Wu, Yufeng Lu, Jiwu Lu et al.USENIX ATC 2024 · 4 citations
- Deep Learning from Imperfectly Labeled Malware DataFahad Alotaibi, Euan Goodbrand, Sergio MaffeisCCS 2025
- Evaluating the Effectiveness and Robustness of Visual Similarity-based Phishing Detection ModelsFujiao Ji, Kiho Lee, Hyungjoon Koo, Wenhao You et al.USENIX Security 2025
Builds on15
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular DomainJinsung Yoon, Yao Zhang, James Jordon, Mihaela van der SchaarNeurIPS 2020 · 370 citations
- Log2vec: A Heterogeneous Graph Embedding Based Approach for Detecting Cyber Threats within EnterpriseFucheng Liu, Yu Wen, Dongxue Zhang, Xihe Jiang et al.CCS 2019 · 314 citations
- Apps, Trackers, Privacy, and Regulators: A Global Study of the Mobile Tracking EcosystemAbbas Razaghpanah, Rishab Nithyanand, Narseo Vallina-Rodriguez, Srikanth Sundaresan et al.NDSS 2018 · 271 citations
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper et al.ICML 2020 · 130 citations
Related papers
- Measuring and Modeling the Label Dynamics of Online Anti-Malware EnginesShuofei Zhu, Jianjun Shi, Limin Yang, Boqin Qin et al.USENIX Security 2020
- Poisoning Self-supervised Learning Based Sequential RecommendationsYanling Wang, Yuchen Liu, Qian Wang, Cong Wang et al.SIGIR 2023 · 16 citations
- UniDoc: Unified Pretraining Framework for Document UnderstandingJiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao et al.NeurIPS 2021 · 118 citations
- Malicious Domain Detection on Out-of-Distribution Gray Data through Graph Contrastive Learning with Structure AggregationHongjie Gu, Daojing He, Xun ZhouKDD 2026
- Fusion Is Not A Simple Ensemble! Towards The Evolving Views in Insider Threat DetectionChengyu Song, Lin Yang, Jianming Zheng, Jingjing Zhang et al.WWW 2026
