When Only Time Will Tell: Interpreting How Transformers Process Local Ambiguities Through the Lens of Restart-Incrementality
Brielen Madureira, Patrick Kahardipraja, David Schlangen
Abstract
Incremental models that process sentences one token at a time will sometimes encounter points where more than one interpretation is possible. Causal models are forced to output one interpretation and continue, whereas models that can revise may edit their previous output as the ambiguity is resolved. In this work, we look at how restart-incremental Transformers build and update internal states, in an effort to shed light on what processes cause revisions not viable in autoregressive models. We propose an interpretable way to analyse the incremental states, showing that their sequential structure encodes information on the garden path effect and its resolution. Our method brings insights on various bidirectional encoders for contextualised meaning representation and dependency parsing, contributing to show their advantage over causal models when it comes to revisions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt et al.ICLR 2024 · 119 citations
- We're Afraid Language Models Aren't Modeling AmbiguityAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr et al.EMNLP 2023 · 35 citations
- Back to the Future: Unsupervised Backprop-based Decoding for Counterfactual and Abductive Commonsense ReasoningLianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula et al.EMNLP 2020 · 25 citations
- Explaining How Transformers Use Context to Build PredictionsJavier Ferrando, Gerard I. Gállego, Ioannis Tsiamas, Marta R. Costa-jussàACL 2023 · 9 citations
Related papers
- Tree Transformer's Disambiguation Ability of Prepositional Phrase Attachment and Garden Path EffectsLingling Zhou, Suzan Verberne, Gijs WijnholdsACL 2024
- Incremental Processing in the Age of Non-Incremental Encoders: An Empirical Assessment of Bidirectional Models for Incremental NLUBrielen Madureira, David SchlangenEMNLP 2020
- Priors in time: Missing inductive biases for language model interpretabilityEkdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur, Valérie Costa et al.ICLR 2026 · 19 citations
- On the Bias of Next-Token Predictors Toward Systematically Inefficient Reasoning: A Shortest-Path Case StudyRiccardo Alberghi, Elizaveta Demyanenko, Luca Biggio, Luca SagliettiNeurIPS 2025 · 2 citations
- When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path CompressionXinnan Dai, Kai Yang, cheng Luo, Shenglai Zeng et al.ICML 2026
