Learning Graph Quantized Tokenizers
Limei Wang, Kaveh Hassani, Si Zhang, Dongqi Fu, Baichuan Yuan, Weilin Cong, Zhigang Hua, Hao Wu, Ning Yao, Bo Long
Abstract
Transformers serve as the backbone architectures of Foundational Models, where domain-specific tokenizers allow them to adapt to various domains. Graph Transformers (GTs) have recently emerged as leading models in geometric deep learning, outperforming Graph Neural Networks (GNNs) in various graph learning tasks. However, the development of tokenizers for graphs has lagged behind other modalities. To address this, we introduce GQT (Graph Quantized Tokenizer), which decouples tokenizer training from Transformer training by leveraging multi-task graph self-supervised learning, yielding robust and generalizable graph tokens. Furthermore, the GQT utilizes Residual Vector Quantization (RVQ) to learn hierarchical discrete tokens, resulting in significantly reduced memory requirements and improved generalization capabilities. By combining the GQT with token modulation, a Transformer encoder achieves state-of-the-art performance on 20 out of 22 benchmarks, including large-scale homophilic and heterophilic datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Towards Effective Federated Graph Foundation Model via Mitigating Knowledge EntanglementYinlin Zhu, Xunkai Li, Jishuo Jia, Miao Hu et al.NeurIPS 2025 · 17 citations
- HeroFilter: Adaptive Spectral Graph Filter for Varying Heterophilic RelationsShuaicheng Zhang, Haohui Wang, Junhong Lin, Xiaojie Guo et al.NeurIPS 2025 · 5 citations
- OwlEye: Zero-Shot Learner for Cross-Domain Graph Data Anomaly DetectionLecheng Zheng, Dongqi Fu, Zihao Li, Jingrui HeICLR 2026 · 2 citations
- Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation LearningZian Zhai, Fan Li, Xingyu Tan, Xiaoyang Wang et al.ICML 2026 · 2 citations
- Structure-Centric Graph Foundation Model via Geometric BasesXiaodong He, Haolan He, Ruiyi Fang, Ming Sun et al.ICML 2026 · 1 citation
Builds on64
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
Related papers
- Pure Transformers are Powerful Graph LearnersJinwoo Kim, Dat Nguyen, Seonwoo Min, Sungjun Cho et al.NeurIPS 2022 · 311 citations
- NAGphormer: A Tokenized Graph Transformer for Node Classification in Large GraphsJinsong Chen, Kaiyuan Gao, Gaichao Li, Kun HeICLR 2023 · 22 citations
- Relational Graph TransformerVijay Prakash Dwivedi, Sri Jaladi, Yangyi Shen, Federico Lopez et al.ICLR 2026 · 35 citations
- Rethinking Tokenizer and Decoder in Masked Graph Modeling for MoleculesZhiyuan Liu, Yaorui Shi, An Zhang, Enzhi Zhang et al.NeurIPS 2023 · 71 citations
- GraphGPT: Generative Pre-trained Graph Eulerian TransformerQifang Zhao, Weidong Ren, Tianyu Li, Hong Liu et al.ICML 2025
