Graph Tokenization for Bridging Graphs and Transformers
Zeyuan Guo, Enmao Diao, Cheng Yang, Chuan Shi
摘要
The success of large pretrained Transformers is closely tied to tokenizers, which convert raw input into discrete symbols. Extending these models to graph-structured data remains a significant challenge. In this work, we introduce a graph tokenization framework that generates sequential representations of graphs by combining reversible graph serialization, which preserves graph information, with Byte Pair Encoding (BPE), a widely adopted tokenizer in large language models (LLMs). To better capture structural information, the graph serialization process is guided by global statistics of graph substructures, ensuring that frequently occurring substructures appear more often in the sequence and can be merged by BPE into meaningful tokens. Empirical results demonstrate that the proposed tokenizer enables Transformers such as BERT to be directly applied to graph benchmarks without architectural modifications. The proposed approach achieves state-of-the-art results on 14 benchmark datasets and frequently outperforms both graph neural networks and specialized graph transformers. This work bridges the gap between graph-structured data and the ecosystem of sequence models. Our code is available at bluehere.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph CompletionQizhuo Xie, Yunhui Liu, Yu Xing, Qianzi Hou 等ACL 2026
- Toward Graph-Tokenizing Large Language Models with Reconstructive Graph Instruction TuningZhongjian Zhang, Xiao Wang, Mengmei Zhang, Jiarui Tan 等WWW 2026
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu 等NeurIPS 2022 · 被引用 1,216 次
- GraphGPT: Graph Instruction Tuning for Large Language ModelsJiabin Tang, Yuhao Yang, Wei Wei, Lei Shi 等SIGIR 2024 · 被引用 182 次
相关 Paper
- Graph Generative Pre-trained TransformerXiaohui Chen, Yinkai Wang, Jiaxing He, Yuanqi Du 等ICML 2025
- Towards Next Graph Token Prediction: Discrete Graph Tokenization for Structural Reasoning in Large Language ModelsZhonghao Wang, Yugang Ji, Zhuonan Zheng, Sheng Zhou 等KDD 2026
- : One LLM Token for Explicit Graph Structural UnderstandingJingyao Wu, Bin Lu, Zijun Di, Xiaoying Gan 等ICLR 2026 · 被引用 2 次
- Learning Graph Quantized TokenizersLimei Wang, Kaveh Hassani, Si Zhang, Dongqi Fu 等ICLR 2025
- From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual ModalitiesWanpeng Zhang, Zilong Xie, Yicheng Feng, Yijiang Li 等ICLR 2025
