Towards Next Graph Token Prediction: Discrete Graph Tokenization for Structural Reasoning in Large Language Models
Zhonghao Wang, Yugang Ji, Zhuonan Zheng, Sheng Zhou, Weigao Wen, Minghao Li, Ming Gu, Zhiyao Zhou, Jiajun Bu
Abstract
Large Language Models (LLMs), empowered by autoregressive next-token prediction, have demonstrated strong reasoning capabilities. Extending this paradigm to graph data requires next graph token prediction, yet existing graph tokens struggle to balance two competing requirements: capturing higher-order structural semantics with high information density, and remaining strictly reversible for faithful decoding. Concretely, first-order textual encodings are reversible but long and semantically sparse, while continuous encodings capture high-level semantics but inevitably lose exact topology. To this end, we propose GraphVulcan, a novel framework that enables discrete reversible and semantic-rich tokenization of graphs using a vocabulary of canonical graphlets. Our approach preserves full structural fidelity while enabling LLMs to natively reason over compositional graphlets through graph-level next-token prediction. We propose a three-stage training paradigm: (1) Structural Semantic Pretraining to learn graph token compositionality, (2) Multi-task Fine-tuning on large-scale CoT-augmented reasoning examples, and (3) Reinforcement Learning to explore and refine reasoning paths. Experiments show that GraphVulcan outperforms first-order encoding baselines across 7 graph reasoning tasks and 3 real-world benchmarks while achieving higher computational efficiency. Our work demonstrates that structure-aware discrete tokenization is a feasible way toward general-purpose graph–language models capable of structural reasoning. Codes are available at: https://github.com/alibaba-behavioral-risk-control/GraphVulcan
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c826d72-d5ea-4aeb-9532-b4252775da6aBuilds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan et al.NeurIPS 2023 · 420 citations
- Talk like a Graph: Encoding Graphs for Large Language ModelsBahare Fatemi, Jonathan Halcrow, Bryan PerozziICLR 2024 · 194 citations
Related papers
- Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language ModelsRunxuan Liu, Xianhao Ou, Xinyan Ma, Jiyuan Wang et al.ACL 2026
- MuSe: Multi-Stage Graph Reasoning via Vision-Language ModelsGuanyu Wang, Xu Chu, Zhijie Tan, Xinrong Chen et al.ACL 2026
- Graph Tokenization for Bridging Graphs and TransformersZeyuan Guo, Enmao Diao, Cheng Yang, Chuan ShiICLR 2026 · 4 citations
- UniGTE: Unified Graph-Text Encoding for Zero-Shot Generalization across Graph Tasks and DomainsDuo Wang, Yuan Zuo, Guangyue Lu, Junjie WuNeurIPS 2025 · 9 citations
- GraphSkill: Documentation-Guided Agentic Hierarchical Retrieval-Augmented Coding for Complex Graph ReasoningFali Wang, Chenglin Weng, Xianren Zhang, Siyuan Hong et al.KDD 2026
