Tokenizing 3D Molecule Structure with Quantized Spherical Coordinates
Kaiyuan Gao, Yusong Wang, Haoxiang Guan, Zun Wang, Qizhi Pei, John E. Hopcroft, Kun He, Lijun Wu
摘要
The application of language models (LMs) to molecular structure generation using line notations such as SMILES and SELFIES has been well-established in the field of cheminformatics. However, extending these models to generate 3D molecular structures presents significant challenges. Two primary obstacles emerge: (1) the difficulty in designing a 3D line notation that ensures SE(3)-invariant atomic coordinates, and (2) the non-trivial task of tokenizing continuous coordinates for use in LMs, which inherently require discrete inputs. To address these challenges, we propose Mol-StrucTok, a novel method for tokenizing 3D molecular structures. Our approach comprises two key innovations: (1) We design a line notation for 3D molecules by extracting local atomic coordinates in a spherical coordinate system. This notation builds upon existing 2D line notations and remains agnostic to their specific forms, ensuring compatibility with various molecular representation schemes. (2) We employ a Vector Quantized Variational Autoencoder (VQ-VAE) to tokenize these coordinates, treating them as generation descriptors. To further enhance the representation, we incorporate neighborhood bond lengths and bond angles as understanding descriptors. Leveraging this tokenization framework, we train a GPT-2 style model for 3D molecular generation tasks. Results demonstrate strong performance with significantly faster generation speeds and competitive chemical stability compared to previous methods. Further, by integrating our learned discrete representations into Graphormer model for property prediction on QM9 dataset, Mol-StrucTok reveals consistent improvements across various molecular properties, underscoring the versatility and robustness of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- E(n) Equivariant Graph Neural NetworksVictor Garcia Satorras, Emiel Hoogeboom, Max WellingICML 2021 · 被引用 1,432 次
- Equivariant Diffusion for Molecule Generation in 3DEmiel Hoogeboom, Victor Garcia Satorras, Clément Vignac, Max WellingICML 2022 · 被引用 865 次
- GraphAF: a Flow-based Autoregressive Model for Molecular Graph GenerationChence Shi, Minkai Xu, Zhaocheng Zhu, Weinan Zhang 等ICLR 2020 · 被引用 532 次
相关 Paper
- Geometry Informed Tokenization of Molecules for Language Model GenerationXiner Li, Limei Wang, Youzhi Luo, Carl Edwards 等ICML 2025
- InertialAR: Autoregressive 3D Molecule Generation with Inertial FramesHaorui Li, weitao du, Yuqiang Li, Hongyu Guo 等ICML 2026 · 被引用 7 次
- Graph VQ-Transformer (GVT): Fast and Accurate Molecular Generation via High-Fidelity Discrete LatentsHaozhuo Zheng, Cheng Wang, Yang LiuAAAI 2026
- NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule GenerationZhiyuan Liu, Yanchen Luo, Han Huang, Enzhi Zhang 等ICLR 2025
- SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAEYongwei Chen, Yushi Lan, Shangchen Zhou, Tengfei Wang 等CVPR 2025
