Can LLMs Fool Graph Learning? Exploring Universal Adversarial Attacks on Text-Attributed Graphs
Zihui Chen, Yuling Wang, Pengfei Jiao, Kai Wu, Xiao Wang, Xiang Ao, Dalin Zhang
Abstract
Text-attributed graphs (TAGs) enhance graph learning by integrating rich textual semantics and topological context for each node. While boosting expressiveness, they also expose new vulnerabilities in graph learning through text-based adversarial surfaces. Recent advances leverage diverse backbones, such as graph neural networks (GNNs) and pre-trained language models (PLMs), to capture both structural and textual information in TAGs. This diversity raises a key question: How can we design universal adversarial attacks that generalize across architectures to assess the security of TAG models? The challenge arises from the stark contrast in how different backbones-GNNs and PLMs-perceive and encode graph patterns, coupled with the fact that many PLMs are only accessible via APIs, limiting attacks to black-box settings. To address this, we propose BadGraph, a novel attack framework that deeply elicits large language models' (LLMs) understanding of general graph knowledge to jointly perturb both node topology and textual semantics. Specifically, we design a target influencer retrieval module that leverages graph priors to construct cross-modally aligned attack shortcuts, thereby enabling efficient LLM-based perturbation reasoning. Experiments show that BadGraph achieves universal and effective attacks across GNN-and LLM-based reasoners, with up to a 76.3% performance drop, while theoretical and empirical analyses confirm its stealthy yet interpretable nature. CCS Concepts • Computing methodologies → Machine learning approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f48a8f02-89af-4e84-afb3-6c9635fd878dBuilds on23
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- One For All: Towards Training One Graph Model For All Classification TasksHao Liu, Jiarui Feng, Lecheng Kong, Ningyue Liang et al.ICLR 2024 · 253 citations
- LLaGA: Large Language and Graph AssistantRunjin Chen, Tong Zhao, Ajay Kumar Jaiswal, Neil Shah et al.ICML 2024 · 180 citations
- Attacking Graph-based Classification via Manipulating the Graph StructureBinghui Wang, Neil Zhenqiang GongCCS 2019 · 175 citations
- Towards More Practical Adversarial Attacks on Graph Neural NetworksJiaqi Ma, Shuangrui Ding, Qiaozhu MeiNeurIPS 2020 · 160 citations
Related papers
- Unveiling the Vulnerability of Graph-LLMs: An Interpretable Multi-Dimensional Adversarial Attack on TAGsBowen Fan, Zhilin Guo, Xunkai Li, Yihan Zhou et al.WWW 2026
- GraphTextack: A Realistic Black-Box Node Injection Attack on LLM-Enhanced GNNsJiaji Ma, Puja Trivedi, Danai KoutraAAAI 2026
- Robustness in Text-Attributed Graph Learning: Insights, Trade-offs, and New DefensesRunlin Lei, Lu Yi, Mingguo He, Pengyu Qiu et al.ICLR 2026 · 1 citation
- Towards Robust Text-Attributed Federated Graph Learning: Multimodal Threats and DefenseZitong Shi, Guancheng Wan, Wenke Huang, Yuxin Wu et al.AAAI 2026
- Generalization Principles for Inference over Text-Attributed Graphs with Large Language ModelsHaoyu Peter Wang, Shikun Liu, Rongzhe Wei, Pan LiICML 2025
