Contrastive losses as generalized models of global epistasis
David H. Brookes, Jakub Otwinowski, Sam Sinai
Abstract
Fitness functions map large combinatorial spaces of biological sequences to properties of interest. Inferring these multimodal functions from experimental data is a central task in modern protein engineering. Global epistasis models are an effective and physically-grounded class of models for estimating fitness functions from observed data. These models assume that a sparse latent function is transformed by a monotonic nonlinearity to emit measurable fitness. Here we demonstrate that minimizing supervised contrastive loss functions, such as the Bradley-Terry loss, is a simple and flexible technique for extracting the sparse latent function implied by global epistasis. We argue by way of a fitness-epistasis uncertainty principle that the nonlinearities in global epistasis models can produce observed fitness functions that do not admit sparse representations, and thus may be inefficient to learn from observations when using a Mean Squared Error (MSE) loss (a common practice). We show that contrastive losses are able to accurately estimate a ranking function from limited data even in regimes where MSE is ineffective and validate the practical utility of this insight by demonstrating that contrastive loss functions result in consistently improved performance on benchmark tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on2
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
- Deep Extrapolation for Attribute-Enhanced GenerationAlvin Chan, Ali Madani, Ben Krause, Nikhil NaikNeurIPS 2021 · 33 citations
Related papers
- Non-identifiability and the Blessings of Misspecification in Models of Molecular FitnessEli N. Weinstein, Alan Nawzad Amin, Jonathan Frazer, Debora S. MarksNeurIPS 2022 · 31 citations
- Rethinking Reward Modeling in Preference-based Large Language Model AlignmentHao Sun, Yunyi Shen, Jean-Francois TonICLR 2025
- Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference ModelYuzhong Hong, Hanshan Zhang, Junwei Bao, Hongfei Jiang et al.ICML 2025
- What Does Preference Learning Recover from Pairwise Comparison Data?Rattana Pukdee, Nina Balcan, Pradeep RavikumarICML 2026 · 1 citation
- Score-Based Density Estimation from Pairwise ComparisonsPetrus Mikkola, Luigi Acerbi, Arto KlamiICLR 2026
