SMI-Editor: Edit-based SMILES Language Model with Fragment-level Supervision
Kangjie Zheng, Siyue Liang, Junwei Yang, Bin Feng, Zequn Liu, Wei Ju, Zhiping Xiao, Ming Zhang
Abstract
SMILES, a crucial textual representation of molecular structures, has garnered significant attention as a foundation for pre-trained language models (LMs). However, most existing pre-trained SMILES LMs focus solely on the single-token level supervision during pre-training, failing to fully leverage the substructural information of molecules. This limitation makes the pre-training task overly simplistic, preventing the models from capturing richer molecular semantic information. Moreover, during pre-training, these SMILES LMs only process corrupted SMILES inputs, never encountering any valid SMILES, which leads to a train-inference mismatch. To address these challenges, we propose SMI-EDITOR, a novel edit-based pre-trained SMILES LM. SMI-EDITOR disrupts substructures within a molecule at random and feeds the resulting SMILES back into the model, which then attempts to restore the original SMILES through an editing process. This approach not only introduces fragment-level training signals, but also enables the use of valid SMILES as inputs, allowing the model to learn how to reconstruct complete molecules from these incomplete structures. As a result, the model demonstrates improved scalability and an enhanced ability to capture fragment-level molecular information. Experimental results show that SMI-EDITOR achieves state-of-the-art performance across multiple downstream molecular tasks, and even outperforming several 3D molecular representation models. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a3f6444-254a-4411-a5ee-ee4af3e10a92Builds on13
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
- InfoGraph: Unsupervised and Semi-supervised Graph-Level Representation Learning via Mutual Information MaximizationFan-Yun Sun, Jordan Hoffmann, Vikas Verma, Jian TangICLR 2020 · 1,010 citations
- Pre-training Molecular Graph Representation with 3D GeometryShengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby et al.ICLR 2022 · 440 citations
- 3D Infomax improves GNNs for Molecular Property PredictionHannes Stärk, Dominique Beaini, Gabriele Corso, Prudencio Tossou et al.ICML 2022 · 269 citations
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng et al.ICLR 2023 · 254 citations
Related papers
- Improving Chemical Understanding of LLMs via SMILES ParsingYunhui Jang, Jaehyung Kim, Sungsoo AhnEMNLP 2025 · 2 citations
- Advancing Molecular Graph-Text Pre-training via Fine-grained AlignmentYibo Li, Yuan Fang, Mengmei Zhang, Chuan ShiKDD 2025
- ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language ModelsKangjie Zheng, Junwei Yang, Siyue Liang, Bin Feng et al.ICML 2025
- MolEditRL: Structure-Preserving Molecular Editing via Discrete Diffusion and Reinforcement LearningYuanxin Zhuang, Dazhong Shen, Ying SunICLR 2026 · 2 citations
- How to Make Large Language Models Generate 100% Valid Molecules?Wen Tao, Jing Tang, Alvin Chan, Bryan Hooi et al.EMNLP 2025
