B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
Luca Zancato, Arjun Seshadri, Yonatan Dukler, Aditya Golatkar, Yantao Shen, Benjamin Bowman, Matthew Trager, Alessandro Achille, Stefano Soatto
摘要
We describe a family of architectures to support transductive inference by allowing memory to grow to a finite but a-priori unknown bound while making efficient use of finite resources for inference. Current architectures use such resources to represent data either eidetically over a finite span ("context"in Transformers), or fading over an infinite span (in State Space Models, or SSMs). Recent hybrid architectures have combined eidetic and fading memory, but with limitations that do not allow the designer or the learning process to seamlessly modulate the two, nor to extend the eidetic memory span. We leverage ideas from Stochastic Realization Theory to develop a class of models called B'MOJO to seamlessly combine eidetic and fading memory within an elementary composable module. The overall architecture can be used to implement models that can access short-term eidetic memory"in-context,"permanent structural memory"in-weights,"fading memory"in-state,"and long-term eidetic memory"in-storage"by natively incorporating retrieval from an asynchronously updated memory. We show that Transformers, existing SSMs such as Mamba, and hybrid architectures such as Jamba are special cases of B'MOJO and describe a basic implementation, to be open sourced, that can be stacked and scaled efficiently in hardware. We test B'MOJO on transductive inference tasks, such as associative recall, where it outperforms existing SSMs and Hybrid models; as a baseline, we test ordinary language modeling where B'MOJO achieves perplexity comparable to similarly-sized Transformers and SSMs up to 1.4B parameters, while being up to 10% faster to train. Finally, we show that B'MOJO's ability to modulate eidetic and fading memory results in better inference on longer sequences tested up to 32K tokens, four-fold the length of the longest sequences seen during training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Titans: Learning to Memorize at Test TimeAli Behrouz, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 被引用 368 次
- DeltaProduct: Improving State-Tracking in Linear RNNs via Householder ProductsJulien Siems, Timur Carstensen, Arber Zela, Frank Hutter 等NeurIPS 2025 · 被引用 75 次
- Distilling to Hybrid Attention Models via KL-Guided Layer SelectionYanhong Li, Songlin Yang, Shawn Tan, Mayank Mishra 等ICLR 2026 · 被引用 17 次
- Gated KalmaNet: A Fading Memory Layer through Test-time Ridge RegressionLiangzu Peng, Aditya Chattopadhyay, Luca Zancato, Elvis Nunez 等CVPR 2026 · 被引用 10 次
- Bridging Expressivity and Scalability with Adaptive Unitary SSMsArjun Karuvally, Franz Nowak, T. Anderson Keller, Carmen Amo Alonso 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
相关 Paper
- Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic ModelsAviv Bick, Kevin Y. Li, Eric P. Xing, J. Zico Kolter 等NeurIPS 2024 · 被引用 78 次
- Jamba: Hybrid Transformer-Mamba Language ModelsBarak Lenz, Opher Lieber, Alan Arazi, Amir Bergman 等ICLR 2025
- Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall CapacityNingyuan Teresa Huang, Miguel Sarabia, Abhinav Moudgil, Pau Rodríguez 等ICML 2025
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingLiliang Ren, Yang Liu, Yadong Lu, Yelong Shen 等ICLR 2025
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language ModelYixing Li, Ruobing Xie, Zhen Yang, Xingwu Sun 等AAAI 2026 · 被引用 3 次
