Removing Biases from Molecular Representations via Information Maximization
Chenyu Wang, Sharut Gupta, Caroline Uhler, Tommi S. Jaakkola
摘要
High-throughput drug screening -- using cell imaging or gene expression measurements as readouts of drug effect -- is a critical tool in biotechnology to assess and understand the relationship between the chemical structure and biological activity of a drug. Since large-scale screens have to be divided into multiple experiments, a key difficulty is dealing with batch effects, which can introduce systematic errors and non-biological associations in the data. We propose InfoCORE, an Information maximization approach for COnfounder REmoval, to effectively deal with batch effects and obtain refined molecular representations. InfoCORE establishes a variational lower bound on the conditional mutual information of the latent representations given a batch identifier. It adaptively reweighs samples to equalize their implied batch distribution. Extensive experiments on drug screening data reveal InfoCORE's superior performance in a multitude of tasks including molecular property prediction and molecule-phenotype retrieval. Additionally, we show results for how InfoCORE offers a versatile framework and resolves general distribution shifts and issues of data fairness by minimizing correlation with spurious features or removing sensitive attributes. The code is available at https://github.com/uhlerlab/InfoCORE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- How Molecules Impact Cells: Unlocking Contrastive PhenoMolecular RetrievalPhilip Fradkin, Puria Azadi Moghadam, Karush Suri, Frederik Wenkel 等NeurIPS 2024 · 被引用 14 次
- Learning Molecular Representation in a CellGang Liu, Srijit Seal, John Arevalo, Zhenwen Liang 等ICLR 2025
- An Information Criterion for Controlled Disentanglement of Multimodal DataChenyu Wang, Sharut Gupta, Xinyi Zhang, Sana Tonekaboni 等ICLR 2025
- Learning Cell-Aware Hierarchical Multi-Modal Representations for Robust Molecular ModelingMengran Li, Zelin Zang, Wenbin Xing, Junzhou Chen 等AAAI 2026
- BiGMINT: Biologically-guided Hierarchical Multimodal Integration for Modeling Multiple Compound Activities in Drug DiscoveryPushpak Pati, Bo Li, Abbas Rayabat Khan, Tomé Albuquerque 等CVPR 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee 等NeurIPS 2020 · 被引用 406 次
- Interpretable and Generalizable Graph Learning via Stochastic Attention MechanismSiqi Miao, Mia Liu, Pan LiICML 2022 · 被引用 288 次
相关 Paper
- Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and FairnessAdam Foster, Árpi Vezér, Craig A. Glastonbury, Páidí Creed 等ICML 2022 · 被引用 7 次
- Variational Learning of Disentangled RepresentationsYuli Slavutsky, Ozgur Beker, David Blei, Bianca DumitrascuICML 2026 · 被引用 1 次
- Efficient Biological Data Acquisition through Inference Set DesignIhor Neporozhnii, Julien Roy, Emmanuel Bengio, Jason S. HartfordICLR 2025
- Asymmetric Contrastive Objectives for Efficient Phenotypic ScreeningLuke Nightingale, Joseph Tuersley, Scott Warchal, Andrea Cairoli 等ICML 2026 · 被引用 1 次
- Learning Cross-Domain Representations for Transferable Drug Perturbations on Single-Cell Transcriptional ResponsesHui Liu, Shikai JinAAAI 2025 · 被引用 1 次
