Attention Projection Mixing with Exogenous Anchors
Jonathan Su
Abstract
Cross-layer reuse of early attention projections can improve optimization and data efficiency, but it creates a structural conflict: the first layer must simultaneously act as a stable, reusable anchor for all deeper layers and as an effective computational block. We demonstrate that this tension constrains the performance of internal-anchor designs. We propose ExoFormer, which resolves the conflict by learning exogenous anchor projections outside the sequential layer stack. We introduce a unified normalized mixing framework that mixes queries, keys, values, and gate logits using learnable coefficients (exploring coefficient granularities: elementwise, headwise, and scalar), and we show that normalizing anchor sources is key to stable reuse. ExoFormer variants consistently outperform their internal-anchor counterparts, and the dynamic variant yields ∼ 1.5 downstream accuracy points while matching validation loss using ∼ 1.5× fewer tokens than Gated Attention. We explain this efficacy via an Offloading Hypothesis: external anchors preserve essential token identity, allowing layers to specialize exclusively in feature transformation. We release code and models to facilitate future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97e20d6a-d056-43f3-a542-b94d898dd009Builds on16
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang et al.NeurIPS 2025 · 336 citations
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
Related papers
- AEA: Adaptive Expert Allocation Improves Sentence Embeddings from Mixture-of-Experts LLMShufan Yang, Zifeng Cheng, Zhiwei Jiang, Qingfeng Qi et al.ACL 2026
- RIFormer: Keep Your Vision Backbone Effective But Removing Token MixerJiahao Wang, Songyang Zhang, Yong Liu, Taiqiang Wu et al.CVPR 2023
- ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable SpecializationAnzhe Cheng, Shukai Duan, Shixuan Li, Chenzhong Yin et al.CVPR 2026 · 8 citations
- Multi-Head Attention as a Source of Catastrophic Forgetting in MoE TransformersAnrui Chen, Ruijun Huang, Xin Zhang, Fang DONG(董方) et al.ICML 2026 · 3 citations
- MixFormer: Mixing Features across Windows and DimensionsQiang Chen, Qiman Wu, Jian Wang, Qinghao Hu et al.CVPR 2022 · 142 citations
