Data-Efficient Molecular Generation with Hierarchical Textual Inversion
Seojin Kim, Jaehyun Nam, Sihyun Yu, Younghoon Shin, Jinwoo Shin
Abstract
Developing an effective molecular generation framework even with a limited number of molecules is often important for its practical deployment, e.g., drug discovery, since acquiring task-related molecular data requires expensive and time-consuming experimental costs. To tackle this issue, we introduce Hierarchical textual Inversion for Molecular generation (HI-Mol), a novel data-efficient molecular generation method. HI-Mol is inspired by the importance of hierarchical information, e.g., both coarseand fine-grained features, in understanding the molecule distribution. We propose to use multilevel embeddings to reflect such hierarchical features based on the adoption of the recent textual inversion technique in the visual domain, which achieves data-efficient image generation. Compared to the conventional textual inversion method in the image domain using a single-level token embedding, our multi-level token embeddings allow the model to effectively learn the underlying low-shot molecule distribution. We then generate molecules based on the interpolation of the multi-level token embeddings. Extensive experiments demonstrate the superiority of HI-Mol with notable data-efficiency. For instance, on QM9, HI-Mol outperforms the prior state-ofthe-art method with 50× less training data. We also show the effectiveness of molecules generated by HI-Mol in low-shot molecular property prediction. Code is available at https: //github.com/Seojin-Kim/HI-Mol .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on27
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
Related papers
- Hierarchical Grammar-Induced Geometry for Data-Efficient Molecular Property PredictionMinghao Guo, Veronika Thost, Samuel W. Song, Adithya Balachandran et al.ICML 2023
- Hierarchical Structure-Property Alignment for Data-Efficient Molecular Generation and EditingZiyu Fan, Zhijian Huang, Yahan Li, Xiaowen Hu et al.AAAI 2026
- NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule GenerationZhiyuan Liu, Yanchen Luo, Han Huang, Enzhi Zhang et al.ICLR 2025
- DiffTMR: Diffusion-based Hierarchical Alignment for Text-Molecule RetrievalChenxu Wang, Dong Zhou, Ting Liu, Jianghao Lin et al.ACM MM 2025
- Text-Guided Molecule Generation with Diffusion Language ModelHaisong Gong, Qiang Liu, Shu Wu, Liang WangAAAI 2024 · 45 citations
