APT-CGLP: Advanced Persistent Threat Hunting via Contrastive Graph-Language Pre-Training
Xuebo Qiu, Mingqi Lv, Yimei Zhang, Tieming Chen, Tiantian Zhu, Qijie Song, Shouling Ji
摘要
Provenance-based threat hunting identifies Advanced Persistent Threats (APTs) on endpoints by correlating attack patterns described in Cyber Threat Intelligence (CTI) with provenance graphs derived from system audit logs. A fundamental challenge in this paradigm lies in the modality gap —the structural and semantic disconnect between provenance graphs and CTI reports. Prior work addresses this by framing threat hunting as a graph matching task: 1) extracting attack graphs from CTI reports, and 2) aligning them with provenance graphs. However, this pipeline incurs severe information loss during graph extraction and demands intensive manual curation, undermining scalability and effectiveness. In this paper, we present APT-CGLP, a novel cross-modal APT hunting system via Contrastive Graph-Language Pre-training, facilitating end-to-end semantic matching between provenance graphs and CTI reports without human intervention. First, empowered by the Large Language Model (LLM), APT-CGLP mitigates data scarcity by synthesizing high-fidelity provenance graph-CTI report pairs, while simultaneously distilling actionable insights from noisy web-sourced CTIs to improve their operational utility. Second, APT-CGLP incorporates a tailored multi-objective training algorithm that synergizes contrastive learning with inter-modal masked modeling, promoting cross-modal attack semantic alignment at both coarse- and fine-grained levels. Extensive experiments on four real-world APT datasets demonstrate that APT-CGLP consistently outperforms state-of-the-art threat hunting baselines in terms of accuracy and efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
相关 Paper
- OCR-APT: Reconstructing APT Stories from Audit Logs using Subgraph Anomaly Detection and LLMsAhmed Aly, Essam Mansour, Amr M. YoussefCCS 2025 · 被引用 2 次
- Enabling Efficient Cyber Threat Hunting With Cyber Threat IntelligencePeng Gao, Fei Shao, Xiaoyuan Liu, Xusheng Xiao 等ICDE 2021 · 被引用 124 次
- POIROT: Aligning Attack Behavior with Kernel Audit Records for Cyber Threat HuntingSadegh M. Milajerdi, Birhanu Eshete, Rigel Gjomemo, V. N. VenkatakrishnanCCS 2019 · 被引用 313 次
- LLMCloudHunter: Harnessing LLMs for Automated Extraction of Detection Rules from Cloud-Based CTIYuval Schwartz, Lavi Ben-Shimol, Dudu Mimran, Yuval Elovici 等WWW 2025 · 被引用 38 次
- KnowHow: Automatically Applying High-Level CTI Knowledge for Interpretable and Accurate Provenance AnalysisYuhan Meng, Shaofei Li, Jiaping Gui, Peng Jiang 等NDSS 2026 · 被引用 8 次
