Rethinking Tokenizer and Decoder in Masked Graph Modeling for Molecules
Zhiyuan Liu, Yaorui Shi, An Zhang, Enzhi Zhang, Kenji Kawaguchi, Xiang Wang, Tat-Seng Chua
摘要
Masked graph modeling excels in the self-supervised representation learning of molecular graphs. Scrutinizing previous studies, we can reveal a common scheme consisting of three key components: (1) graph tokenizer, which breaks a molecular graph into smaller fragments (i.e., subgraphs) and converts them into tokens; (2) graph masking, which corrupts the graph with masks; (3) graph autoencoder, which first applies an encoder on the masked graph to generate the representations, and then employs a decoder on the representations to recover the tokens of the original graph. However, the previous MGM studies focus extensively on graph masking and encoder, while there is limited understanding of tokenizer and decoder. To bridge the gap, we first summarize popular molecule tokenizers at the granularity of node, edge, motif, and Graph Neural Networks (GNNs), and then examine their roles as the MGM's reconstruction targets. Further, we explore the potential of adopting an expressive decoder in MGM. Our results show that a subgraph-level tokenizer and a sufficiently expressive decoder with remask decoding have a large impact on the encoder's representation learning. Finally, we propose a novel MGM method SimSGT, featuring a Simple GNN-based Tokenizer (SGT) and an effective decoding strategy. We empirically validate that our method outperforms the existing molecule self-supervised learning methods. Our codes and checkpoints are available at https://github.com/syr-cn/SimSGT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- EXGC: Bridging Efficiency and Explainability in Graph CondensationJunfeng Fang, Xinglin Li, Yongduo Sui, Yuan Gao 等WWW 2024 · 被引用 30 次
- Pre-Training Graph Neural Networks on Molecules by Using Subgraph-Conditioned Graph Information BottleneckVan Thuy Hoang, O-Joun LeeAAAI 2025 · 被引用 19 次
- MolParser: End-to-End Visual Recognition of Molecule Structures in the WildXi Fang, Jiankun Wang, Xiaochen Cai, Shangqian Chen 等ICCV 2025 · 被引用 8 次
- Hi-GMAE: Hierarchical Graph Masked AutoencodersChuang Liu, Zelin Yao, Xueqi Ma, Mukun Chen 等WWW 2026 · 被引用 3 次
- Learning 3D Anisotropic Noise Distributions Improves Molecular Force FieldsXixian Liu, Rui Jiao, Zhiyuan Liu, Yurou Liu 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper38
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
相关 Paper
- 3D-GSRD: 3D Molecular Graph Auto-Encoder with Selective Re-mask DecodingChang Wu, Zhiyuan Liu, Wen Shu, Liang Wang 等NeurIPS 2025 · 被引用 1 次
- What's Behind the Mask: Understanding Masked Graph Modeling for Graph AutoencodersJintang Li, Ruofan Wu, Wangbin Sun, Liang Chen 等KDD 2023 · 被引用 89 次
- GraphMAE: Self-Supervised Masked Graph AutoencodersZhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong 等KDD 2022 · 被引用 533 次
- KPGT: Knowledge-Guided Pre-training of Graph Transformer for Molecular Property PredictionHan Li, Dan Zhao, Jianyang ZengKDD 2022 · 被引用 55 次
- Self-supervised Masked Graph Autoencoder via Structure-aware CurriculumHaoyang Li, Xin Wang, Zeyang Zhang, Zongyuan Wu 等ICML 2025
