CFD: Learning Generalized Molecular Representation via Concept-Enhanced Feedback Disentanglement
Aming Wu, Cheng Deng
Abstract
To accelerate biochemical research, e.g., drug and protein discovery, molecular representation learning (MRL) has attracted much attention. However, most existing methods follow the closed-set assumption that training and testing data share identical distribution, which limits their generalization abilities in out-of-distribution (OOD) cases. In this paper, we explore designing a new disentangled mechanism for learning generalized molecular representation that exhibits robustness against distribution shifts. And an approach of Concept-Enhanced Feedback Disentanglement (CFD) is proposed, whose goal is to exploit the feedback mechanism to learn distribution-agnostic representation. Specifically, we first propose two dedicated variational encoders to separately decompose distribution-agnostic and spurious features. Then, a set of molecule-aware concepts are tapped to focus on invariant substructure characteristics. By fusing these concepts into the disentangled distribution-agnostic features, the generalization ability of the learned molecular representation could be further enhanced. Next, we execute iteratively the disentangled operations based on a feedback received from the previous output. Finally, based on the outputs of multiple feedback iterations, we construct a self-supervised objective to promote the variational encoders to possess the disentangled capability.
In the experiments, our method is verified on multiple real-world molecular datasets. The significant performance gains over state-of-the-art baselines demonstrate that our method can effectively disentangle generalized molecular representation in the presence of various distribution shifts. The source code will be released at https://github.com/AmingWu/MoleculeCFD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54e34adf-ad6c-4a85-8db4-7066d58cea74Cited by top-tier papers1
Ask how each one uses itBuilds on32
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu et al.NeurIPS 2022 · 1,216 citations
Related papers
- Learning Substructure Invariance for Out-of-Distribution Molecular RepresentationsNianzu Yang, Kaipeng Zeng, Qitian Wu, Xiaosong Jia et al.NeurIPS 2022 · 133 citations
- Learning Invariant Molecular Representation in Latent Discrete SpaceXiang Zhuang, Qiang Zhang, Keyan Ding, Yatao Bian et al.NeurIPS 2023 · 41 citations
- Disentangled Graph LLM for Molecule Graph Editing under Distribution ShiftsYang Yao, Xin Wang, Yuan Meng, Zeyang Zhang et al.WWW 2026
- Invariant Conditional Molecular Generation Under Distribution ShiftChunyu Hu, Tianyin Liao, Yicheng Sui, Ran Zhang et al.AAAI 2026
- C-Disentanglement: Discovering Causally-Independent Generative Factors under an Inductive Bias of ConfounderXiaoyu Liu, Jiaxin Yuan, Bang An, Yuancheng Xu et al.NeurIPS 2023 · 13 citations
