Enhancing Activity Prediction Models in Drug Discovery with the Ability to Understand Human Language
Philipp Seidl, Andreu Vall, Sepp Hochreiter, Günter Klambauer
Abstract
Activity and property prediction models are the central workhorses in drug discovery and materials sciences, but currently they have to be trained or fine-tuned for new tasks. Without training or fine-tuning, scientific language models could be used for such low-data tasks through their announced zero- and few-shot capabilities. However, their predictive quality at activity prediction is lacking. In this work, we envision a novel type of activity prediction model that is able to adapt to new prediction tasks at inference time, via understanding textual information describing the task. To this end, we propose a new architecture with separate modules for chemical and natural language inputs, and a contrastive pre-training objective on data from large biochemical databases. In extensive experiments, we show that our method CLAMP yields improved predictive performance on few-shot learning benchmarks and zero-shot problems in drug discovery. We attribute the advances of our method to the modularized architecture and to our pre-training objective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd7d82b0-b0a4-4641-a822-61996357effeCited by top-tier papers14
- GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot LearningHaiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu et al.NeurIPS 2023 · 97 citations
- MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal AdapterZhiyuan Liu, Sihang Li, Yanchen Luo, Hao Fei et al.EMNLP 2023 · 34 citations
- CHEMREASONER: Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical FeedbackHenry W. Sprueill, Carl Edwards, Khushbu Agarwal, Mariefel V. Olarte et al.ICML 2024 · 19 citations
- MolTailor: Tailoring Chemical Molecular Representation to Specific Tasks via Text PromptsHaoqiang Guo, Sendong Zhao, Haochun Wang, Yanrui Du et al.AAAI 2024 · 17 citations
- Learning Multi-view Molecular Representations with Structured and Unstructured KnowledgeYizhen Luo, Kai Yang, Massimo Hong, Xing Yi Liu et al.KDD 2024 · 9 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
Related papers
- CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal PoseXu Zhang, Wen Wang, Zhe Chen, Yufei Xu et al.CVPR 2023
- CrystalICL: Enabling In-Context Learning for Crystal GenerationRuobing Wang, Qiaoyu Tan, Yili Wang, Ying Wang et al.EMNLP 2025 · 3 citations
- Tag-LLM: Repurposing General-Purpose LLMs for Specialized DomainsJunhong Shen, Neil A. Tenenholtz, James Brian Hall, David Alvarez-Melis et al.ICML 2024 · 60 citations
- Unifying Molecular and Textual Representations via Multi-task Language ModellingDimitrios Christofidellis, Giorgio Giannone, Jannis Born, Ole Winther et al.ICML 2023 · 126 citations
- Generalizable Drug-Target Interaction Prediction via ESM-2 Representations and Progressive Contrastive Curriculum LearningQianyang Wu, Jingwei Lv, Zilong Zhang, Feifei CuiAAAI 2026
