MolTRES: Improving Chemical Language Representation Learning for Molecular Property Prediction
Jun-Hyung Park, Yeachan Kim, Mingyu Lee, Hyuntae Park, SangKeun Lee
Abstract
Chemical representation learning has gained increasing interest due to the limited availability of supervised data in fields such as drug and materials design.This interest particularly extends to chemical language representation learning, which involves pre-training Transformers on SMILES sequences -textual descriptors of molecules.Despite its success in molecular property prediction, current practices often lead to overfitting and limited scalability due to early convergence.In this paper, we introduce a novel chemical language representation learning framework, called MolTRES, to address these issues.MolTRES incorporates generator-discriminator training, allowing the model to learn from more challenging examples that require structural understanding.In addition, we enrich molecular representations by transferring knowledge from scientific literature by integrating external materials embedding.Experimental results show that our models outperform existing state-of-the-art models on popular molecular property prediction tasks.github.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e90ecf11-af84-4e51-ac53-f0709049fdb8Cited by top-tier papers1
Ask how each one uses itBuilds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
Related papers
- MolTailor: Tailoring Chemical Molecular Representation to Specific Tasks via Text PromptsHaoqiang Guo, Sendong Zhao, Haochun Wang, Yanrui Du et al.AAAI 2024 · 17 citations
- KPGT: Knowledge-Guided Pre-training of Graph Transformer for Molecular Property PredictionHan Li, Dan Zhao, Jianyang ZengKDD 2022 · 55 citations
- Translation between Molecules and Natural LanguageCarl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke et al.EMNLP 2022 · 112 citations
- RTMol: Rethinking Molecule-text Alignment in a Round-trip ViewLetian Chen, Runhan Shi, Gufeng Yu, Yang YangAAAI 2026 · 1 citation
- Uni-Mol: A Universal 3D Molecular Representation Learning FrameworkGengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng et al.ICLR 2023 · 254 citations
