LAIT: Efficient Multi-Segment Encoding in Transformers with Layer-Adjustable Interaction
Jeremiah Milbauer, Annie Louis, Mohammad Javad Hosseini, Alex Fabrikant, Donald Metzler, Tal Schuster
Abstract
Transformer encoders contextualize token representations by attending to all other tokens at each layer, leading to quadratic increase in compute effort with the input length. In practice, however, the input text of many NLP tasks can be seen as a sequence of related segments (e.g., the sequence of sentences within a passage, or the hypothesis and premise in NLI). While attending across these segments is highly beneficial for many tasks, we hypothesize that this interaction can be delayed until later encoding stages. To this end, we introduce Layer-Adjustable Interactions in Transformers (LAIT). Within LAIT, segmented inputs are first encoded independently, and then jointly. This partial two-tower architecture bridges the gap between a Dual Encoder's ability to precompute representations for segments and a fully self-attentive Transformer's capacity to model cross-segment attention. The LAIT framework effectively leverages existing pretrained Transformers and converts them into the hybrid of the two aforementioned architectures, allowing for easy and intuitive control over the performance-efficiency tradeoff. Experimenting on a wide range of NLP tasks, we find LAIT able to reduce 30-50% of the attention FLOPs on many tasks, while preserving high accuracy; in some practical settings, LAIT could reduce actual latency by orders of magnitude.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73e8e0e8-71ef-456f-8f73-d1638fb0fce6Cited by top-tier papers2
- Pre-computed memory or on-the-fly encoding? A hybrid approach to retrieval augmentation makes the most of your computeMichiel de Jong, Yury Zemlyanskiy, Nicholas FitzGerald, Joshua Ainslie et al.ICML 2023 · 20 citations
- Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRASangmin Bae, Adam Fisch, Hrayr Harutyunyan, Ziwei Ji et al.ICLR 2025
Builds on13
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
Related papers
- Native Hybrid Attention for Efficient Sequence ModelingJusen Du, Jiaxi Hu, Zhang Tao, Weigao Sun et al.ACL 2026 · 8 citations
- Learn-to-Share: A Hardware-friendly Transfer Learning Framework Exploiting Computation and Parameter SharingCheng Fu, Hanxian Huang, Xinyun Chen, Yuandong Tian et al.ICML 2021 · 28 citations
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingZihang Dai, Guokun Lai, Yiming Yang, Quoc LeNeurIPS 2020 · 273 citations
- One Model, Many Budgets: Elastic Latent Interfaces for Diffusion TransformersMoayed Haji Ali, Willi Menapace, Ivan Skorokhodov, Dogyun Park et al.CVPR 2026 · 4 citations
- Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention ModelsDifan Deng, Andreas B. Winje, Lukas Fehring, Marius LindauerICML 2026
