SALSA: Semantically-Aware Latent Space Autoencoder
Kathryn E. Kirchoff, Travis Maxfield, Alexander Tropsha, Shawn M. Gomez
Abstract
In deep learning for drug discovery, chemical data are often represented as simplified molecular-input line-entry system (SMILES) sequences which allow for straightforward implementation of natural language processing methodologies, one being the sequence-to-sequence autoencoder. However, we observe that training an autoencoder solely on SMILES is insufficient to learn molecular representations that are semantically meaningful, where semantics are defined by the structural (graph-to-graph) similarities between molecules. We demonstrate by example that autoencoders may map structurally similar molecules to distant codes, resulting in an incoherent latent space that does not respect the structural similarities between molecules. To address this shortcoming we propose Semantically-Aware Latent Space Autoencoder (SALSA), a transformer-autoencoder modified with a contrastive task, tailored specifically to learn graph-to-graph similarity between molecules. Formally, the contrastive objective is to map structurally similar molecules (separated by a single graph edit) to nearby codes in the latent space. To accomplish this, we generate a novel dataset comprised of sets of structurally similar molecules and opt for a supervised contrastive loss that is able to incorporate full sets of positive samples. We compare SALSA to its ablated counterparts, and show empirically that the composed training objective (reconstruction and contrastive task) leads to a higher quality latent space that is more 1) structurally-aware, 2) semantically continuous, and 3) property-aware.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8516ab7c-d8ef-49a1-a0ed-bc8381685bc2Builds on4
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Self-supervised Graph-level Representation Learning with Local and Global StructureMinghao Xu, Hang Wang, Bingbing Ni, Hongyu Guo et al.ICML 2021 · 248 citations
- Optimus: Organizing Sentences via Pre-trained Modeling of a Latent SpaceChunyuan Li, Xiang Gao, Yuan Li, Baolin Peng et al.EMNLP 2020 · 132 citations
- Educating Text Autoencoders: Latent Representation Guidance via DenoisingTianxiao Shen, Jonas Mueller, Regina Barzilay, Tommi S. JaakkolaICML 2020 · 74 citations
Related papers
- Path-Aware and Structure-Preserving Generation of Synthetically Accessible MoleculesJuhwan Noh, Dae-Woong Jeong, Kiyoung Kim, Sehui Han et al.ICML 2022 · 11 citations
- Variational Graph Autoencoding as Cheap Supervision for AMR Coreference ResolutionIrene Li, Linfeng Song, Kun Xu, Dong YuACL 2022 · 12 citations
- GraphMAE: Self-Supervised Masked Graph AutoencodersZhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong et al.KDD 2022 · 533 citations
- S²Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual ScreeningBowei He, Bowen Gao, Yankai Chen, Yanyan Lan et al.AAAI 2026 · 1 citation
- 3DLinker: An E(3) Equivariant Variational Autoencoder for Molecular Linker DesignYinan Huang, Xingang Peng, Jianzhu Ma, Muhan ZhangICML 2022 · 65 citations
