Chain and Causal Attention for Efficient Entity Tracking
Erwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre Allauzen
Abstract
This paper investigates the limitations of transformers for entity-tracking tasks in large language models. We identify a theoretical constraint, showing that transformers require at least log 2 (n + 1) layers to handle entity tracking with n state changes. To address this issue, we propose an efficient and frugal enhancement to the standard attention mechanism, enabling it to manage long-term dependencies more efficiently. By considering attention as an adjacency matrix, our model can track entity states with a single layer. Empirical results demonstrate significant improvements in entity tracking datasets while keeping competitive performance on standard natural language modeling. Our modified attention allows us to achieve the same performance with drastically fewer layers. Additionally, our enhanced mechanism reveals structured internal representations of attention. Extensive experiments on both toy and complex datasets validate our approach. Our contributions include theoretical insights, an improved attention mechanism, and empirical validation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e84308fc-118a-4d9f-b580-1754cd204ad9Cited by top-tier papers4
- PaTH Attention: Position Encoding via Accumulating Householder TransformationsSonglin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan et al.NeurIPS 2025 · 36 citations
- MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning ModelsVanya Cohen, Ray MooneyICML 2026 · 2 citations
- Trading Complexity for Expressivity Through Structured Generalized Linear Token MixingErwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre AllauzenICML 2026 · 1 citation
- Language models can learn implicit multi-hop reasoning, but only if they have lots of training dataYuekun Yao, Yupei Du, Dawei Zhu, Michael Hahn et al.EMNLP 2025
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- Do Language Models Track Entities Across State Changes?Zilu Tang, Qiao Zhao, Gabriel Franco, Derry Wijaya et al.ICML 2026
- Low-Rank Bottleneck in Multi-head Attention ModelsSrinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J. Reddi et al.ICML 2020 · 130 citations
- Evolving Attention with Residual ConvolutionsYujing Wang, Yaming Yang, Jiangang Bai, Mingliang Zhang et al.ICML 2021 · 43 citations
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingNikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov et al.ICLR 2024 · 113 citations
- Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language ModelsZeping Yu, Yonatan Belinkov, Sophia AnaniadouEMNLP 2025 · 2 citations
