A Structured Observation Distribution for Generative Biological Sequence Prediction and Forecasting
Eli N. Weinstein, Debora S. Marks
Abstract
Generative probabilistic modeling of biological sequences has widespread existing and potential application across biology and biomedicine, from evolutionary biology to epidemiology to protein design. Many standard sequence analysis methods preprocess data using a multiple sequence alignment (MSA) algorithm, one of the most widely used computational methods in all of science. However, as we show in this article, training generative probabilistic models with MSA preprocessing leads to statistical pathologies in the context of sequence prediction and forecasting. To address these problems, we propose a principled drop-in alternative to MSA preprocessing in the form of a structured observation distribution (the ``MuE" distribution). The MuE is a latent alignment model in which not only the alignment variable but also the regressor sequence can be latent. We prove theoretically that the MuE distribution comprehensively generalizes popular methods for inferring biological sequence alignments, and provide a precise characterization of how such biological models have differed from natural language latent alignment models. We show empirically that models that use the MuE as an observation distribution outperform comparable methods across a variety of datasets, and apply MuE models to a novel problem for generative probabilistic sequence models: forecasting pathogen evolution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4c9c2b5-e653-4ccc-a40e-b6b8f5f2793aCited by top-tier papers4
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
- Non-identifiability and the Blessings of Misspecification in Models of Molecular FitnessEli N. Weinstein, Alan Nawzad Amin, Jonathan Frazer, Debora S. MarksNeurIPS 2022 · 31 citations
- A generative nonparametric Bayesian model for whole genomesAlan Nawzad Amin, Eli N. Weinstein, Debora S. MarksNeurIPS 2021 · 9 citations
- A Kernelized Stein Discrepancy for Biological SequencesAlan Nawzad Amin, Eli N. Weinstein, Debora Susan MarksICML 2023 · 4 citations
Related papers
- Evolution-Inspired Loss Functions for Protein Representation LearningChengyue Gong, Adam R. Klivans, James Loy, Tianlong Chen et al.ICML 2024 · 10 citations
- RMSAGen: Integrating Multiple Sequence Alignment for Function RNA DesignJiyue Jiang, Yanyu Chen, Qingchuan Zhang, Jiayi Li et al.AAAI 2026 · 1 citation
- Markovian Linguistic-Temporal Bridge: Unlocking the Potential of LLMs for Time Series ForecastingSiming Sun, Kai Zhang, Xuejun Jiang, Wenchao Meng et al.ACL 2026
- Retrieved Sequence Augmentation for Protein Representation LearningChang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin et al.EMNLP 2024 · 4 citations
- Multiple sequence alignment as a sequence-to-sequence learning problemEdo Dotan, Yonatan Belinkov, Oren Avram, Elya Wygoda et al.ICLR 2023 · 34 citations
