Towards Next Graph Token Prediction: Discrete Graph Tokenization for Structural Reasoning in Large Language Models
Zhonghao Wang, Yugang Ji, Zhuonan Zheng, Sheng Zhou, Weigao Wen, Minghao Li, Ming Gu, Zhiyao Zhou, Jiajun Bu
摘要
Large Language Models (LLMs), empowered by autoregressive next-token prediction, have demonstrated strong reasoning capabilities. Extending this paradigm to graph data requires next graph token prediction, yet existing graph tokens struggle to balance two competing requirements: capturing higher-order structural semantics with high information density, and remaining strictly reversible for faithful decoding. Concretely, first-order textual encodings are reversible but long and semantically sparse, while continuous encodings capture high-level semantics but inevitably lose exact topology. To this end, we propose GraphVulcan, a novel framework that enables discrete reversible and semantic-rich tokenization of graphs using a vocabulary of canonical graphlets. Our approach preserves full structural fidelity while enabling LLMs to natively reason over compositional graphlets through graph-level next-token prediction. We propose a three-stage training paradigm: (1) Structural Semantic Pretraining to learn graph token compositionality, (2) Multi-task Fine-tuning on large-scale CoT-augmented reasoning examples, and (3) Reinforcement Learning to explore and refine reasoning paths. Experiments show that GraphVulcan outperforms first-order encoding baselines across 7 graph reasoning tasks and 3 real-world benchmarks while achieving higher computational efficiency. Our work demonstrates that structure-aware discrete tokenization is a feasible way toward general-purpose graph–language models capable of structural reasoning. Codes are available at: https://github.com/alibaba-behavioral-risk-control/GraphVulcan
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan 等NeurIPS 2023 · 被引用 420 次
- Talk like a Graph: Encoding Graphs for Large Language ModelsBahare Fatemi, Jonathan Halcrow, Bryan PerozziICLR 2024 · 被引用 194 次
相关 Paper
- Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language ModelsRunxuan Liu, Xianhao Ou, Xinyan Ma, Jiyuan Wang 等ACL 2026
- MuSe: Multi-Stage Graph Reasoning via Vision-Language ModelsGuanyu Wang, Xu Chu, Zhijie Tan, Xinrong Chen 等ACL 2026
- Graph Tokenization for Bridging Graphs and TransformersZeyuan Guo, Enmao Diao, Cheng Yang, Chuan ShiICLR 2026 · 被引用 4 次
- UniGTE: Unified Graph-Text Encoding for Zero-Shot Generalization across Graph Tasks and DomainsDuo Wang, Yuan Zuo, Guangyue Lu, Junjie WuNeurIPS 2025 · 被引用 9 次
- GraphSkill: Documentation-Guided Agentic Hierarchical Retrieval-Augmented Coding for Complex Graph ReasoningFali Wang, Chenglin Weng, Xianren Zhang, Siyuan Hong 等KDD 2026
