APT-CGLP: Advanced Persistent Threat Hunting via Contrastive Graph-Language Pre-Training
Xuebo Qiu, Mingqi Lv, Yimei Zhang, Tieming Chen, Tiantian Zhu, Qijie Song, Shouling Ji
Abstract
Provenance-based threat hunting identifies Advanced Persistent Threats (APTs) on endpoints by correlating attack patterns described in Cyber Threat Intelligence (CTI) with provenance graphs derived from system audit logs. A fundamental challenge in this paradigm lies in the modality gap —the structural and semantic disconnect between provenance graphs and CTI reports. Prior work addresses this by framing threat hunting as a graph matching task: 1) extracting attack graphs from CTI reports, and 2) aligning them with provenance graphs. However, this pipeline incurs severe information loss during graph extraction and demands intensive manual curation, undermining scalability and effectiveness. In this paper, we present APT-CGLP, a novel cross-modal APT hunting system via Contrastive Graph-Language Pre-training, facilitating end-to-end semantic matching between provenance graphs and CTI reports without human intervention. First, empowered by the Large Language Model (LLM), APT-CGLP mitigates data scarcity by synthesizing high-fidelity provenance graph-CTI report pairs, while simultaneously distilling actionable insights from noisy web-sourced CTIs to improve their operational utility. Second, APT-CGLP incorporates a tailored multi-objective training algorithm that synergizes contrastive learning with inter-modal masked modeling, promoting cross-modal attack semantic alignment at both coarse- and fine-grained levels. Extensive experiments on four real-world APT datasets demonstrate that APT-CGLP consistently outperforms state-of-the-art threat hunting baselines in terms of accuracy and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- OCR-APT: Reconstructing APT Stories from Audit Logs using Subgraph Anomaly Detection and LLMsAhmed Aly, Essam Mansour, Amr M. YoussefCCS 2025 · 2 citations
- Enabling Efficient Cyber Threat Hunting With Cyber Threat IntelligencePeng Gao, Fei Shao, Xiaoyuan Liu, Xusheng Xiao et al.ICDE 2021 · 124 citations
- POIROT: Aligning Attack Behavior with Kernel Audit Records for Cyber Threat HuntingSadegh M. Milajerdi, Birhanu Eshete, Rigel Gjomemo, V. N. VenkatakrishnanCCS 2019 · 313 citations
- LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTIYuval Schwartz, Lavi Ben-Shimol, Dudu Mimran, Yuval Elovici et al.WWW 2025 · 38 citations
- KnowHow: Automatically Applying High-Level CTI Knowledge for Interpretable and Accurate Provenance AnalysisYuhan Meng, Shaofei Li, Jiaping Gui, Peng Jiang et al.NDSS 2026 · 8 citations
