A new framework for evaluating model out-of-distribution generalisation for the biochemical domain
Raúl Fernández-Díaz, Hoang Thanh Lam, Vanessa López, Denis C. Shields
摘要
A bstract Quantifying model generalization to out-of-distribution data has been a longstanding challenge in machine learning. Addressing this issue is crucial for leveraging machine learning in scientific discovery, where models must generalize to new molecules or materials. Current methods typically split data into train and test sets using various criteria — temporal, sequence identity, scaffold, or random cross-validation — before evaluating model performance. However, with so many splitting criteria available, existing approaches offer limited guidance on selecting the most appropriate one, and they do not provide mechanisms for incorporating prior knowledge about the target deployment distribution(s). To tackle this problem, we have developed a novel metric, AU-GOOD, which quantifies expected model performance under conditions of increasing dissimilarity between train and test sets, while also accounting for prior knowledge about the target deployment distribution(s), when available. This metric is broadly applicable to biochemical entities, including proteins, small molecules, nucleic acids, or cells; as long as a relevant similarity function is defined for them. Recognizing the wide range of similarity functions used in biochemistry, we propose criteria to guide the selection of the most appropriate metric for partitioning. We also introduce a new partitioning algorithm that generates more challenging test sets, and we propose statistical methods for comparing models based on AU-GOOD. Finally, we demonstrate the insights that can be gained from this framework by applying it to two different use cases: developing predictors for pharmaceutical properties of small molecules, and using protein language models as embeddings to build biophysical property predictors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Rethinking the generalization of drug target affinity prediction algorithms via similarity aware evaluationChenbin Zhang, Zhiqiang Hu, Chuchu Jiang, Wen Chen 等ICLR 2025
- Exploring Chemical Space with Score-based Out-of-distribution GenerationSeul Lee, Jaehyeong Jo, Sung Ju HwangICML 2023 · 被引用 110 次
- Learning Substructure Invariance for Out-of-Distribution Molecular RepresentationsNianzu Yang, Kaipeng Zeng, Qitian Wu, Xiaosong Jia 等NeurIPS 2022 · 被引用 133 次
- DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery - a Focus on Affinity Prediction Problems with Noise AnnotationsYuanfeng Ji, Lu Zhang, Jiaxiang Wu, Bingzhe Wu 等AAAI 2023 · 被引用 65 次
- Distributional Priors Guided Diffusion for Generating 3D Molecules in Low Data RegimesHaokai Hong, Wanyu Lin, Ming Yang, Kay Chen TanAAAI 2026 · 被引用 1 次
