FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications
Kieran Didi, Sarah Alamdari, Alex Lu, Bruce Wittmann, Kadina Johnston, Ava Amini, Ali Madani, Maya Czeneszew, Christian Dallago, Kevin Yang
Abstract
Machine learning methods that predict protein fitness from sequence remain sensitive to changes in data distributions, limiting generalization across common conditions encountered in protein engineering. Practically, protein engineers are thus left wondering about the effective utility of ML tools. The FLIP benchmark established protocols for testing generalization under some domain shifts, but it was limited to measurements of thermostability, binding, and viral capsid viability. We introduce FLIP2, a protein fitness benchmark spanning seven new datasets, including enzymes, protein-protein interactions, and light-sensitive proteins, as well as splits that measure generalization relevant to real-world protein engineering campaigns. Evaluating a suite of benchmark models across these datasets and splits reveals that simpler models often matched or outperformed fine-tuned protein language models on FLIP2, challenging the utility of existing transfer learning techniques. Provenance for all datasets has been recorded and we redistribute all data CC-BY 4.0 to facilitate continued progress.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 813f63e5-41dc-4b82-a98b-51648632c1edBuilds on4
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang et al.ICML 2024 · 61 citations
- Steering Generative Models with Experimental Data for Protein Fitness OptimizationJason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo et al.NeurIPS 2025 · 13 citations
Related papers
- Metalic: Meta-Learning In-Context with Protein Language ModelsJacob Beck, Shikha Surana, Manus McAuliffe, Oliver Bent et al.ICLR 2025 · 3 citations
- RankFlow: Property-aware Transport for Protein OptimizationLu Yu, Wei Xiang, Kang Han, Gaowen Liu et al.ICLR 2026
- One protein is all you needAnton Bushuiev, Roman Bushuiev, Olga Pimenova, Nikola Zadorozhny et al.ICLR 2026 · 1 citation
- A new framework for evaluating model out-of-distribution generalisation for the biochemical domainRaúl Fernández-Díaz, Hoang Thanh Lam, Vanessa López, Denis C. ShieldsICLR 2025 · 5 citations
- Learning to engineer protein flexibilityPetr Kouba, Joan Planas-Iglesias, Jirí Damborský, Jirí Sedlár et al.ICLR 2025
