Efficient Code Analysis via Graph Representation Learning-Guided Large Language Models
Hang Gao, Tao Peng, Baoquan Cui, Hong Huang, Fengge Wu, Zhao Junsuo, Jian Zhang
Abstract
Large Language Models (LLMs) have significantly advanced code analysis tasks, yet they struggle to detect malicious behaviors fragmented across files, whose intricate dependencies easily get lost in the vast amount of benign code. We therefore propose a graph-centric attention acquisition pipeline that enhances LLMs' ability to localize malicious behavior. The approach parses a project into a code graph, uses an LLM to encode nodes with semantic and structural signals, and trains a Graph Neural Network (GNN) under sparse supervision. The GNN performs an initial detection, and by interpreting these predictions, identifies key code sections that are most likely to contain malicious behavior. These influential regions are then used to guide the LLM's attention for in-depth analysis. This strategy significantly reduces interference from irrelevant context while maintaining low annotation costs. Extensive experiments show that the method consistently outperforms existing approaches on multiple public and custom datasets, highlighting its potential for practical deployment in software security scenarios. Codes can be found in https://github.com/Epiphaniespt/GMLLM.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0da8b3f3-69c9-4716-8284-a019c583e3daBuilds on17
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu et al.ICSE 2024 · 264 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
- Large Language Models for Code Analysis: Do LLMs Really Do Their Job?Chongzhou Fang, Ning Miao, Shaurya Srivastav, Jialin Liu et al.USENIX Security 2024 · 110 citations
- DONAPI: Malicious NPM Packages Detector using Behavior Sequence Knowledge MappingCheng Huang, Nannan Wang, Ziyan Wang, Siqi Sun et al.USENIX Security 2024 · 38 citations
Related papers
- G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent SystemsShilong Wang, Guibin Zhang, Miao Yu, Guancheng Wan et al.ACL 2025 · 37 citations
- MalGuard: Towards Real-Time, Accurate, and Actionable Detection of Malicious Packages in PyPI EcosystemXingan Gao, Xiaobing Sun, Sicong Cao, Kaifeng Huang et al.USENIX Security 2025
- Are LLM-Enhanced Graph Neural Networks Robust Against Poisoning Attacks?Yuhang Ma, Jie Wang, Zheng YanS&P 2026 · 4 citations
- Maltracker: A Fine-Grained NPM Malware Tracker Copiloted by LLM-Enhanced DatasetZeliang Yu, Ming Wen, Xiaochen Guo, Hai JinISSTA 2024 · 16 citations
- GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language ModelYunhe Pang, Bo Chen, Fanjin Zhang, Yanghui Rao et al.KDD 2025
