Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models
Tobias Grantner, Emanuel Sallinger, Martin Flechl
摘要
Transformer-based embedding models suffer from quadratic computational and linear memory complexity, limiting their utility for long sequences. We propose recurrent architectures as an efficient alternative, introducing a vertically chunked inference strategy that enables fast embedding generation with memory usage that becomes constant in the input length once it exceeds the vertical chunk size. By fine-tuning Mamba2 models, we demonstrate their viability as general-purpose text embedders, achieving competitive performance across a range of benchmarks while maintaining a substantially smaller memory footprint compared to transformer-based counterparts. We empirically validate the applicability of our inference strategy to Mamba2, RWKV, and xLSTM models, confirming consistent runtime-memory trade-offs across architectures and establishing recurrent models as a compelling alternative to transformers for efficient embedding generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
相关 Paper
- Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory AccessXiang Hu, Jiaqi Leng, Jun Zhao, Kewei Tu 等NeurIPS 2025 · 被引用 7 次
- DiffuMamba: High-Throughput Diffusion LMs with Mamba BackboneVaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël, Eugene Belilovsky 等ICML 2026
- Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context UnderstandingKai Liu, Zhan Su, Peijie Dong, Fengran Mo 等ICLR 2026 · 被引用 3 次
- MesaNet: Sequence Modeling by Locally Optimal Test-Time TrainingJohannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari 等ICLR 2026 · 被引用 45 次
- RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence ModelingXiuying Wei, Anunay Yadav, Razvan Pascanu, Caglar GulcehreNeurIPS 2025 · 被引用 3 次
