PRSformer: Disease Prediction from Million-Scale Individual Genotypes
Payam Dibaeinia, Chris German, Suyash Shringarpure, Adam Auton, Aly A. Khan
摘要
Predicting disease risk from DNA presents an unprecedented emerging challenge as biobanks approach population scale sizes (N > 10 6 individuals) with ultra-highdimensional features (L > 10 5 genotypes). Current methods, often linear and reliant on summary statistics, fail to capture complex genetic interactions and discard valuable individual-level information. We introduce PRSformer, a scalable deep learning architecture designed for end-to-end, multitask disease prediction directly from million-scale individual genotypes. PRSformer employs neighborhood attention, achieving linear O(L) complexity per layer, making Transformers tractable for genome-scale inputs. Crucially, PRSformer utilizes a stacking of these efficient attention layers, progressively increasing the effective receptive field to model local dependencies (e.g., within linkage disequilibrium blocks) before integrating information across wider genomic regions. This design, tailored for genomics, allows PRSformer to learn complex, potentially non-linear and long-range interactions directly from raw genotypes. We demonstrate PRSformer's effectiveness using a unique large private cohort (N ≈ 5M) for predicting 18 autoimmune and inflammatory conditions using L ≈ 140k variants. PRSformer significantly outperforms highly optimized linear models trained on the same individual-level data and state-of-the-art summary-statistic-based methods (LDPred2) derived from the same cohort, quantifying the benefits of non-linear modeling and multitask learning at scale. Furthermore, experiments reveal that the advantage of non-linearity emerges primarily at large sample sizes (N > 1M), and that a multi-ancestry trained model improves generalization, establishing PRSformer as a new framework for deep learning in population-scale genomics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
- Faster Neighborhood Attention: Reducing the O(n^2) Cost of Self Attention at the Threadblock LevelAli Hassani, Wen-Mei Hwu, Humphrey ShiNeurIPS 2024 · 被引用 23 次
相关 Paper
- Training Flexible Models of Genetic Variant Effects from Functional Annotations using Accelerated Linear AlgebraAlan Nawzad Amin, Andres Potapczynski, Andrew Gordon WilsonICML 2025
- Neuroformer: Multimodal and Multitask Generative Pretraining for Brain DataAntonis Antoniades, Yiyi Yu, Joseph Canzano, William Yang Wang 等ICLR 2024 · 被引用 21 次
- Neural Pharmacodynamic State Space ModelingZeshan M. Hussain, Rahul G. Krishnan, David A. SontagICML 2021 · 被引用 12 次
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan 等AAAI 2021 · 被引用 675 次
- Proxy-Bridged Game Transformer for Interactive Extreme Motion PredictionYanwen Fang, Wenqi Jia, Xu Cao, Peng-Tao Jiang 等ICCV 2025
