A new framework for evaluating model out-of-distribution generalisation for the biochemical domain
Raúl Fernández-Díaz, Hoang Thanh Lam, Vanessa López, Denis C. Shields
Abstract
A bstract Quantifying model generalization to out-of-distribution data has been a longstanding challenge in machine learning. Addressing this issue is crucial for leveraging machine learning in scientific discovery, where models must generalize to new molecules or materials. Current methods typically split data into train and test sets using various criteria — temporal, sequence identity, scaffold, or random cross-validation — before evaluating model performance. However, with so many splitting criteria available, existing approaches offer limited guidance on selecting the most appropriate one, and they do not provide mechanisms for incorporating prior knowledge about the target deployment distribution(s). To tackle this problem, we have developed a novel metric, AU-GOOD, which quantifies expected model performance under conditions of increasing dissimilarity between train and test sets, while also accounting for prior knowledge about the target deployment distribution(s), when available. This metric is broadly applicable to biochemical entities, including proteins, small molecules, nucleic acids, or cells; as long as a relevant similarity function is defined for them. Recognizing the wide range of similarity functions used in biochemistry, we propose criteria to guide the selection of the most appropriate metric for partitioning. We also introduce a new partitioning algorithm that generates more challenging test sets, and we propose statistical methods for comparing models based on AU-GOOD. Finally, we demonstrate the insights that can be gained from this framework by applying it to two different use cases: developing predictors for pharmaceutical properties of small molecules, and using protein language models as embeddings to build biophysical property predictors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e572d0e6-851c-4a1f-8c46-ece3780a3e1cBuilds on1
Related papers
- Rethinking the generalization of drug target affinity prediction algorithms via similarity aware evaluationChenbin Zhang, Zhiqiang Hu, Chuchu Jiang, Wen Chen et al.ICLR 2025
- Exploring Chemical Space with Score-based Out-of-distribution GenerationSeul Lee, Jaehyeong Jo, Sung Ju HwangICML 2023 · 110 citations
- Learning Substructure Invariance for Out-of-Distribution Molecular RepresentationsNianzu Yang, Kaipeng Zeng, Qitian Wu, Xiaosong Jia et al.NeurIPS 2022 · 133 citations
- DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery - a Focus on Affinity Prediction Problems with Noise AnnotationsYuanfeng Ji, Lu Zhang, Jiaxiang Wu, Bingzhe Wu et al.AAAI 2023 · 65 citations
- Distributional Priors Guided Diffusion for Generating 3D Molecules in Low Data RegimesHaokai Hong, Wanyu Lin, Ming Yang, Kay Chen TanAAAI 2026 · 1 citation
