Do Transformers Parse while Predicting the Masked Word?
Haoyu Zhao, Abhishek Panigrahi, Rong Ge, Sanjeev Arora
Abstract
Pre-trained language models have been shown to encode linguistic structures like parse trees in their embeddings while being trained unsupervised. Some doubts have been raised whether the models are doing parsing or only some computation weakly correlated with it. Concretely: (a) Is it possible to explicitly describe transformers with realistic embedding dimensions, number of heads, etc. that are capable of doing parsing -or even approximate parsing? (b) Why do pre-trained models capture parsing structure? This paper takes a step toward answering these questions in the context of generative modeling with PCFGs. We show that masked language models like BERT or RoBERTa of moderate sizes can approximately execute the Inside-Outside algorithm for the English PCFG (Marcus et al., 1993) . We also show that the Inside-Outside algorithm is optimal for masked language modeling loss on the PCFG-generated data. We conduct probing experiments on models pre-trained on PCFG-generated data to show that this not only allows recovery of approximate parse tree, but also recovers marginal span probabilities computed by the Inside-Outside algorithm, which suggests an implicit bias of masked language modeling towards this algorithm. * When | P| < c| Ĩ|, we can simulate the computations in the final layer using c layers with | Ĩ| heads instead of | P| heads. Additionally, we can decrease the embedding size by only storing probabilities for relevant non-terminals.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dad9110f-126b-48a8-9be2-5fade92970f0Cited by top-tier papers13
- Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammarsKaiyue Wen, Yuchen Li, Bingbin Liu, Andrej RisteskiNeurIPS 2023 · 32 citations
- Sample Complexity and Representation Ability of Test-time Scaling ParadigmsBaihe Huang, Shanda Li, Tianhao Wu, Yiming Yang et al.ICLR 2026 · 11 citations
- Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical GuidelinesYuchen Li, Alexandre Kirchmeyer, Aashay Mehta, Yilong Qin et al.ICML 2024 · 5 citations
- Deep networks learn to parse uniform-depth context-free languages from local statisticsJack T. Parley, Francesco Cagnetta, Matthieu WyartICML 2026 · 4 citations
- Context-free Recognition with TransformersSelim Jerad, Anej Svete, Sophie Hao, Ryan Cotterell et al.ICML 2026 · 3 citations
Builds on15
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz et al.NeurIPS 2022 · 267 citations
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 158 citations
- Inductive Biases and Variable Creation in Self-Attention MechanismsBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Cyril ZhangICML 2022 · 154 citations
- Probing Natural Language Inference Models through Semantic FragmentsKyle Richardson, Hai Hu, Lawrence S. Moss, Ashish SabharwalAAAI 2020 · 152 citations
- Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with TransformersColin Wei, Yining Chen, Tengyu MaNeurIPS 2022 · 117 citations
Related papers
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
- Contextual Distortion Reveals Constituency: Masked Language Models are Implicit ParsersJiaxi Li, Wei LuACL 2023 · 3 citations
- Unveiling the Black Box of PLMs with Semantic Anchors: Towards Interpretable Neural Semantic ParsingLunyiu Nie, Jiuding Sun, Yanlin Wang, Lun Du et al.AAAI 2023 · 9 citations
- Efficient Constituency Parsing by PointingThanh-Tung Nguyen, Xuan-Phi Nguyen, Shafiq R. Joty, Xiaoli LiACL 2020 · 11 citations
- Evaluating the Impact of Model Scale for Compositional Generalization in Semantic ParsingLinlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi et al.EMNLP 2022 · 21 citations
