DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery - a Focus on Affinity Prediction Problems with Noise Annotations
Yuanfeng Ji, Lu Zhang, Jiaxiang Wu, Bingzhe Wu, Lanqing Li, Long-Kai Huang, Tingyang Xu, Yu Rong, Jie Ren, Ding Xue, Houtim Lai, Wei Liu
摘要
AI-aided drug discovery (AIDD) is gaining popularity due to its potential to make the search for new pharmaceuticals faster, less expensive, and more effective. Despite its extensive use in numerous fields (e.g., ADMET prediction, virtual screening), little research has been conducted on the out-of-distribution (OOD) learning problem with noise. We present DrugOOD, a systematic OOD dataset curator and benchmark for AIDD. Particularly, we focus on the drug-target binding affinity prediction problem, which involves both macromolecule (protein target) and small-molecule (drug compound). DrugOOD offers an automated dataset curator with user-friendly customization scripts, rich domain annotations aligned with biochemistry knowledge, realistic noise level annotations, and rigorous benchmarking of SOTA OOD algorithms, as opposed to only providing fixed datasets. Since the molecular data is often modeled as irregular graphs using graph neural network (GNN) backbones, DrugOOD also serves as a valuable testbed for graph OOD learning problems. Extensive empirical studies have revealed a significant performance gap between in-distribution and out-of-distribution experiments, emphasizing the need for the development of more effective schemes that permit OOD generalization under noise for AIDD. * Equal contribution. Order was determined by tossing a coin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- On the Stability of Expressive Positional Encodings for GraphsYinan Huang, William Lu, Joshua Robinson, Yu Yang 等ICLR 2024 · 被引用 32 次
- Pairwise Alignment Improves Graph Domain AdaptationShikun Liu, Deyu Zou, Han Zhao, Pan LiICML 2024 · 被引用 27 次
- Context-Guided Diffusion for Out-of-Distribution Molecular and Protein DesignLeo Klarner, Tim G. J. Rudner, Garrett M. Morris, Charlotte M. Deane 等ICML 2024 · 被引用 18 次
- Optimizing OOD Detection in Molecular Graphs: A Novel Approach with Diffusion ModelsXu Shen, Yili Wang, Kaixiong Zhou, Shirui Pan 等KDD 2024 · 被引用 12 次
- FedGOG: Federated Graph Out-of-Distribution Generalization with Diffusion Data Exploration and Latent Embedding DecorrelationPengyang Zhou, Chaochao Chen, Weiming Liu, Xinting Liao 等AAAI 2025 · 被引用 8 次
它引用的顶会 Paper11
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang 等ICML 2021 · 被引用 1,163 次
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie 等NeurIPS 2020 · 被引用 1,113 次
相关 Paper
- Improving Out-of-Distribution Generalization in Graphs via Hierarchical Semantic EnvironmentsYinhua Piao, Sangseon Lee, Yijingxiu Lu, Sun KimCVPR 2024 · 被引用 6 次
- Learning Causally Invariant Representations for Out-of-Distribution Generalization on GraphsYongqiang Chen, Yonggang Zhang, Yatao Bian, Han Yang 等NeurIPS 2022 · 被引用 246 次
- Learning Substructure Invariance for Out-of-Distribution Molecular RepresentationsNianzu Yang, Kaipeng Zeng, Qitian Wu, Xiaosong Jia 等NeurIPS 2022 · 被引用 133 次
- A new framework for evaluating model out-of-distribution generalisation for the biochemical domainRaúl Fernández-Díaz, Hoang Thanh Lam, Vanessa López, Denis C. ShieldsICLR 2025 · 被引用 5 次
- Knowledge Enhanced Representation Learning for Drug DiscoveryThanh Lam Hoang, Marco Luca Sbodio, Marcos Martínez Galindo, Mykhaylo Zayats 等AAAI 2024 · 被引用 9 次
