Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven Priors
Ido Amos, Jonathan Berant, Ankit Gupta
Abstract
Modeling long-range dependencies across sequences is a longstanding goal in machine learning and has led to architectures, such as state space models, that dramatically outperform Transformers on long sequences. However, these impressive empirical gains have been by and large demonstrated on benchmarks (e.g. Long Range Arena), where models are randomly initialized and trained to predict a target label from an input sequence. In this work, we show that random initialization leads to gross overestimation of the differences between architectures and that pretraining with standard denoising objectives, using only the downstream task data, leads to dramatic gains across multiple architectures and to very small gaps between Transformers and state space models (SSMs). In stark contrast to prior works, we find vanilla Transformers to match the performance of S4 on Long Range Arena when properly pretrained, and we improve the best reported results of SSMs on the PathX-256 task by 20 absolute points. Subsequently, we analyze the utility of previously-proposed structured parameterizations for SSMs and show they become mostly redundant in the presence of data-driven initialization obtained through pretraining. Our work shows that, when evaluating different architectures on supervised tasks, incorporation of data-driven priors via pretraining is essential for reliable performance estimation, and can be done efficiently.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c25960b9-2ec9-4ca4-bc4e-10afe02c45b8Cited by top-tier papers20
- LLM-ESR: Large Language Models Enhancement for Long-tailed Sequential RecommendationQidong Liu, Xian Wu, Yejing Wang, Zijian Zhang et al.NeurIPS 2024 · 154 citations
- Single Image Unlearning: Efficient Machine Unlearning in Multimodal Large Language ModelsJiaqi Li, Qianshan Wei, Chuanyi Zhang, Guilin Qi et al.NeurIPS 2024 · 62 citations
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang et al.NeurIPS 2024 · 50 citations
- Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex EigenvaluesAntonio Orvieto, Soham De, Caglar Gulcehre, Razvan Pascanu et al.ICML 2024 · 36 citations
- Perceiving Longer Sequences With Bi-Directional Cross-Attention TransformersMarkus Hiller, Krista A. Ehinger, Tom DrummondNeurIPS 2024 · 23 citations
Builds on23
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- Viewing Transformers Through the Lens of Long Convolutions LayersItamar Zimerman, Lior WolfICML 2024 · 4 citations
- Tuning Frequency Bias of State Space ModelsAnnan Yu, Dongwei Lyu, Soon Hoe Lim, Michael W. Mahoney et al.ICLR 2025
- On the Parameterization and Initialization of Diagonal State Space ModelsAlbert Gu, Karan Goel, Ankit Gupta, Christopher RéNeurIPS 2022 · 690 citations
- Simplified State Space Layers for Sequence ModelingJimmy T. H. Smith, Andrew Warrington, Scott W. LindermanICLR 2023 · 78 citations
- Simple Hardware-Efficient Long Convolutions for Sequence ModelingDaniel Y. Fu, Elliot L. Epstein, Eric Nguyen, Armin W. Thomas et al.ICML 2023 · 72 citations
