GAGE: Genetic Algorithm-Based Graph Explainer for Malware Analysis
Mohd Saqib, Benjamin C. M. Fung, Philippe Charland, Andrew Walenstein
Abstract
Malware analysts often prefer reverse engineering using Call Graphs, Control Flow Graphs (CFGs), and Data Flow Graphs (DFGs), which involves the utilization of black-box Deep Learning (DL) models. The proposed research introduces a structured pipeline for reverse engineering-based analysis, offering promising results compared to state-of-the-art methods and providing high-level interpretability for malicious code blocks in subgraphs. We propose the Canonical Executable Graph (CEG) as a new representation of Portable Executable (PE) files, uniquely incorporating syntactical and semantic information into its node embeddings. At the same time, edge features capture structural aspects of PE files. This is the first work to present a PE file representation encompassing syntactical, semantic, and structural characteristics, whereas previous efforts typically focused solely on syntactic or structural properties. Furthermore, recognizing the limitations of existing graph explanation methods within Explainable Artificial Intelligence (XAI) for malware analysis, primarily due to the specificity of malicious files, we introduce Genetic Algorithm-based Graph Explainer (GAGE). GAGE operates on the CEG, striving to identify a precise subgraph relevant to predicted malware families. Through experiments and comparisons, our proposed pipeline exhibits substantial improvements in model robustness scores and discriminative power compared to the previous benchmarks. Furthermore, we have successfully used GAGE in practical applications on real-world data, producing meaningful insights and interpretability. This research offers a robust solution to enhance cybersecurity by delivering a transparent and accurate understanding of malware behaviour. Moreover, the proposed algorithm is specialized in handling graph-based data, effectively dissecting complex content and isolating influential nodes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b13a74b0-5c7f-4cea-b5b3-2e0a2f6d1688Related papers
- MalGraph: Hierarchical Graph Neural Networks for Robust Windows Malware DetectionXiang Ling, Lingfei Wu, Wei Deng, Zhenqing Qu et al.INFOCOM 2022 · 47 citations
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su et al.CCS 2018 · 336 citations
- API2Vec: Learning Representations of API Sequences for Malware DetectionLei Cui, Jiancong Cui, Yuede Ji, Zhiyu Hao et al.ISSTA 2023 · 37 citations
- EMC: A Semantic-Enhanced Malware Classification Method with Robustness and ScalabilityHaojun Zhao, Yueming Wu, Zhen Li, Deqing ZouICSE 2026
- DeepReflect: Discovering Malicious Functionality through Binary ReconstructionEvan Downing, Yisroel Mirsky, Kyuhong Park, Wenke LeeUSENIX Security 2021 · 43 citations
