Multi-View Empowered Structural Graph Wordification for Language Models
Zipeng Liu, Likang Wu, Ming He, Zhong Guan, Hongke Zhao, Nan Feng
Abstract
Significant efforts have been dedicated to integrating the powerful Large Language Models (LLMs) with diverse modalities, particularly focusing on the fusion of language, vision and audio data. However, the graph-structured data, which is inherently rich in structural and domain-specific knowledge, has not yet been gracefully adapted to LLMs. Existing methods either describe the graph with raw text, suffering the loss of graph structural information, or feed Graph Neural Network (GNN) embeddings into LLMs at the cost of losing explainable prompt semantics. To bridge this gap, we introduce an end-to-end modality-aligning framework for LLM-graph alignment: Dual-Residual Vector Quantized-Variational AutoEncoder, namely Dr.E. Our approach is purposefully designed to facilitate token-level alignment with LLMs, enabling an effective translation of the intrinsic 'language' of graphs into comprehensible natural language. We also manage to enhance LLMs' more robust structural understanding of graphs by incorporating multiple views of the central nodes based on their surrounding nodes at various distances. Our experimental evaluations on standard graph tasks demonstrate competitive performance against other state-ofthe-art (SOTA) approaches. Additionally, our framework ensures certain visual interpretability, efficiency, and robustness, marking the promising successful endeavor to achieve token-level alignment between LLMs and GNNs. Our code is
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 723bee5e-0ab7-4b65-b8b4-55b5b1e9eefbCited by top-tier papers3
- : One LLM Token for Explicit Graph Structural UnderstandingJingyao Wu, Bin Lu, Zijun Di, Xiaoying Gan et al.ICLR 2026 · 2 citations
- Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured DataZhong Guan, Likang Wu, Hongke Zhao, Ming He et al.ICML 2025
- DIN: Dual Impulse Network for Multi-view Representation LearningYilin Wu, Weihong Lin, Renjie Lin, Zihan Fang et al.AAAI 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- NodeFormer: A Scalable Graph Structure Learning Transformer for Node ClassificationQitian Wu, Wentao Zhao, Zenan Li, David P. Wipf et al.NeurIPS 2022 · 472 citations
- GraphFormers: GNN-nested Transformers for Representation Learning on Textual GraphJunhan Yang, Zheng Liu, Shitao Xiao, Chaozhuo Li et al.NeurIPS 2021 · 262 citations
Related papers
- ReaLM: Residual Quantization Bridges Knowledge Graph Embeddings and Large Language ModelsWenbin Guo, Xin Wang, Jiaoyan Chen, Lingbing Guo et al.WWW 2026
- ERAlign: Energy-based Representation Alignment of GNNs and LLMs on Text-attributed GraphsXianlin Zeng, Fan Xia, Xiangyu ChenICML 2026
- GALLa: Graph Aligned Large Language Models for Improved Source Code UnderstandingZiyin Zhang, Hang Yu, Sage Lee, Peng Di et al.ACL 2025 · 11 citations
- Killing Two Birds with One Stone: Cross-modal Reinforced Prompting for Graph and Language TasksWenyuan Jiang, Wenwei Wu, Le Zhang, Zixuan Yuan et al.KDD 2024 · 4 citations
- GITA: Graph to Visual and Textual Integration for Vision-Language Graph ReasoningYanbin Wei, Shuai Fu, Weisen Jiang, Zejian Zhang et al.NeurIPS 2024 · 56 citations
