Fast and Low-Cost Genomic Foundation Models via Outlier Removal
Haozheng Luo, Chenghao Qiu, Maojiang Su, Zhihan Zhou, Zoe Mehta, Guo Ye, Jerry Yao-Chieh Hu, Han Liu
摘要
To address the challenge of scarce computational resources in genomic modeling, we introduce GERM, a genomic foundation model with strong compression performance and fast adaptability. GERM improves upon models like DNABERT-2 by eliminating outliers that hinder low-rank adaptation and post-training quantization, enhancing both efficiency and robustness. We replace the vanilla attention layer with an outlier-free mechanism inspired by associative memory models. By removing outliers during both pre-training and fine-tuning, this approach accelerates adaptation, reduces computational costs, and enhances quantization robustness within acceptable loss margins. Additionally, we propose GERM-T, a strategy that employs small-step continual learning within the outlier-free framework, leveraging original checkpoints to avoid retraining from scratch. Empirically, GERM improves fine-tuning performance by 37.98% and quantization by 64.34% over the baseline model. It also reduces average kurtosis by 92.14% and maximum infinity norm by 82.77%. Compared to leading methods, GERM consistently delivers superior performance, offering a practical solution for genomic modeling in resource-constrained settings. Code is available at https://github.com/MAGICS-LAB/GERM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Tokenization to Transfer: Do Genomic Foundation Models Learn Good Representations?Kirill Vishniakov, Karthik Viswanathan, Aleksandr Medvedev, Praveenkumar Kanithi 等ICLR 2026 · 被引用 16 次
- RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human MobilityHaoyu He, Haozheng Luo, Yan Chen, Qi (Cheems) WangNeurIPS 2025 · 被引用 7 次
- FROST: Filtering Reasoning Outliers with Attention for Efficient ReasoningHaozheng Luo, Zhuolin Jiang, Md Zahid Hasan, Yan Chen 等ICLR 2026 · 被引用 4 次
- Why Do Some Inputs Break Low-Bit LLM Quantization?Ting-Yun Chang, Muru Zhang, Jesse Thomason, Robin JiaEMNLP 2025 · 被引用 1 次
- On the Relationship Between Activation Outliers and Feature Death in Sparse AutoencodersElana Simon, Etowah Adams, James ZouICML 2026
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu 等ICLR 2024 · 被引用 395 次
相关 Paper
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation ModelsWeimin Wu, Xuefeng Song, Yibo Wen, Qinjie Lin 等ICML 2026
- Outlier-Efficient Hopfield Layers for Large Transformer-Based ModelsJerry Yao-Chieh Hu, Pei-Hsuan Chang, Haozheng Luo, Hong-Yu Chen 等ICML 2024 · 被引用 46 次
- dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence LearningArnav Shah, Junzhe Li, Parsa Idehpour, Adibvafa Fallahpour 等ICML 2026
- Genomics Data Lossless Compression with (S, K)-Mer Encoding and Deep Neural NetworksHui Sun, Liping Yi, Huidong Ma, Yongxia Sun 等AAAI 2025 · 被引用 2 次
- Quantized Gradient Projection for Memory-Efficient Continual LearningDongjun Kim, Seohyeon Cha, Huancheng Chen, Chaining Wang 等ICLR 2026
