InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames
Haorui Li, weitao du, Yuqiang Li, Hongyu Guo, Shengchao Liu
Abstract
Transformer-based autoregressive models have emerged as a unifying paradigm across modalities such as text and images, but their extension to 3D molecule generation remains underexplored. The gap stems from two fundamental challenges: (1) how to tokenize molecules into a canonical 1D sequence of tokens that is invariant to both SE(3) transformations and atom index permutations, and (2) how to design an architecture capable of modeling hybrid atom-based tokens that couple discrete atom types with continuous 3D coordinates. To address these challenges, we introduce InertialAR. It first performs generation-oriented canonical tokenization by aligning each molecule to a canonical inertial frame and reordering atoms, thereby converting arbitrary 3D structures into a unique, SE(3)- and permutation-invariant sequence of tokens for autoregressive generation. Built upon this canonical tokenization, we propose geometric positional encoding (GeoPE), which endows Transformer attention with 3D geometric awareness. Finally, InertialAR utilizes a hierarchical autoregressive paradigm to decode the next atom, consecutively predicting the atom type and 3D coordinates via Diffusion Loss. Experimentally, InertialAR achieves state-of-the-art performance on 8 of the 10 evaluation metrics for unconditional generation across QM9, GEOM-Drugs, and B3LYP. Moreover, it significantly outperforms baselines in controllable generation for targeted chemical functionality, attaining state-of-the-art results across all 5 metrics. Code is available at github.com/HaoruiLi46/InertialAR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9bc7323f-8fed-4cde-8d1c-c9e9af273916Cited by top-tier papers2
- A Resolution-Agnostic Geometric Transformer for Chromosome Modeling Using Inertial FrameYize Zhou, Haorui Li, Shengchao LiuICLR 2026 · 2 citations
- Rigidity-Aware Geometric Pretraining for Protein Design and Conformational EnsemblesZhanghan Ni, Yanjing Li, Zeju Qiu, Bernhard Schölkopf et al.ICLR 2026 · 2 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- E(n) Equivariant Graph Neural NetworksVictor Garcia Satorras, Emiel Hoogeboom, Max WellingICML 2021 · 1,432 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
Related papers
- Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion ModelingYanchen Luo, Zhiyuan Liu, Yi Zhao, Sihang Li et al.NeurIPS 2025 · 11 citations
- Geometric Transformer with Interatomic Positional EncodingYusong Wang, Shaoning Li, Tong Wang, Bin Shao et al.NeurIPS 2023 · 25 citations
- Geometry Informed Tokenization of Molecules for Language Model GenerationXiner Li, Limei Wang, Youzhi Luo, Carl Edwards et al.ICML 2025
- Tokenizing 3D Molecule Structure with Quantized Spherical CoordinatesKaiyuan Gao, Yusong Wang, Haoxiang Guan, Zun Wang et al.KDD 2026 · 5 citations
- Sampling 3D Molecular Conformers with Diffusion TransformersJ. Thorben Frank, Winfried Ripken, Gregor Lied, Klaus-Robert Müller et al.NeurIPS 2025 · 7 citations
