Non-identifiability and the Blessings of Misspecification in Models of Molecular Fitness
Eli N. Weinstein, Alan Nawzad Amin, Jonathan Frazer, Debora S. Marks
摘要
Understanding the consequences of mutation for molecular fitness and function is a fundamental problem in biology. Recently, generative probabilistic models have emerged as a powerful tool for estimating fitness from evolutionary sequence data, with accuracy sufficient to predict both laboratory measurements of function and disease risk in humans, and to design novel functional proteins. Existing techniques rest on an assumed relationship between density estimation and fitness estimation, a relationship that we interrogate in this article. We prove that fitness is not identifiable from observational sequence data alone, placing fundamental limits on our ability to disentangle fitness landscapes from phylogenetic history. We show on real datasets that perfect density estimation in the limit of infinite data would, with high confidence, result in poor fitness estimation; current models perform accurate fitness estimation because of, not despite, misspecification. Our results challenge the conventional wisdom that bigger models trained on bigger datasets will inevitably lead to better fitness estimation, and suggest novel estimation strategies going forward.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- PoET: A generative model of protein families as sequences-of-sequencesTimothy F. Truong Jr., Tristan BeplerNeurIPS 2023 · 被引用 96 次
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of ProteinsAadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran 等NeurIPS 2025 · 被引用 63 次
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang 等ICML 2024 · 被引用 61 次
- Understanding protein function with a multimodal retrieval-augmented foundation modelTimothy F. Truong Jr., Tristan BeplerNeurIPS 2025 · 被引用 15 次
- From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language ModelsCharles W. J. Pugh, Paulina G. Nuñez-Valencia, Mafalda Dias, Jonathan FrazerNeurIPS 2025 · 被引用 12 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu 等NeurIPS 2021 · 被引用 969 次
- MSA TransformerRoshan Rao, Jason Liu, Robert Verkuil, Joshua Meier 等ICML 2021 · 被引用 686 次
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 被引用 260 次
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado 等ICML 2022 · 被引用 236 次
相关 Paper
- Protein Language Model Fitness is a Matter of PreferenceCade W. Gordon, Amy X. Lu, Pieter AbbeelICLR 2025
- A Variational Perspective on Generative Protein Fitness OptimizationLea Bogensperger, Dominik Narnhofer, Ahmed Allam, Konrad Schindler 等ICML 2025
- Steering Generative Models with Experimental Data for Protein Fitness OptimizationJason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo 等NeurIPS 2025 · 被引用 13 次
- Contrastive losses as generalized models of global epistasisDavid H. Brookes, Jakub Otwinowski, Sam SinaiNeurIPS 2024 · 被引用 8 次
- A Structured Observation Distribution for Generative Biological Sequence Prediction and ForecastingEli N. Weinstein, Debora S. MarksICML 2021 · 被引用 12 次
