Do Transformers Parse while Predicting the Masked Word?
Haoyu Zhao, Abhishek Panigrahi, Rong Ge, Sanjeev Arora
摘要
Pre-trained language models have been shown to encode linguistic structures like parse trees in their embeddings while being trained unsupervised. Some doubts have been raised whether the models are doing parsing or only some computation weakly correlated with it. Concretely: (a) Is it possible to explicitly describe transformers with realistic embedding dimensions, number of heads, etc. that are capable of doing parsing -or even approximate parsing? (b) Why do pre-trained models capture parsing structure? This paper takes a step toward answering these questions in the context of generative modeling with PCFGs. We show that masked language models like BERT or RoBERTa of moderate sizes can approximately execute the Inside-Outside algorithm for the English PCFG (Marcus et al., 1993) . We also show that the Inside-Outside algorithm is optimal for masked language modeling loss on the PCFG-generated data. We conduct probing experiments on models pre-trained on PCFG-generated data to show that this not only allows recovery of approximate parse tree, but also recovers marginal span probabilities computed by the Inside-Outside algorithm, which suggests an implicit bias of masked language modeling towards this algorithm. * When | P| < c| Ĩ|, we can simulate the computations in the final layer using c layers with | Ĩ| heads instead of | P| heads. Additionally, we can decrease the embedding size by only storing probabilities for relevant non-terminals.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammarsKaiyue Wen, Yuchen Li, Bingbin Liu, Andrej RisteskiNeurIPS 2023 · 被引用 32 次
- Sample Complexity and Representation Ability of Test-time Scaling ParadigmsBaihe Huang, Shanda Li, Tianhao Wu, Yiming Yang 等ICLR 2026 · 被引用 11 次
- Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical GuidelinesYuchen Li, Alexandre Kirchmeyer, Aashay Mehta, Yilong Qin 等ICML 2024 · 被引用 5 次
- Deep networks learn to parse uniform-depth context-free languages from local statisticsJack T. Parley, Francesco Cagnetta, Matthieu WyartICML 2026 · 被引用 4 次
- Context-free Recognition with TransformersSelim Jerad, Anej Svete, Sophie Hao, Ryan Cotterell 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper15
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz 等NeurIPS 2022 · 被引用 267 次
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 被引用 158 次
- Inductive Biases and Variable Creation in Self-Attention MechanismsBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Cyril ZhangICML 2022 · 被引用 154 次
- Probing Natural Language Inference Models through Semantic FragmentsKyle Richardson, Hai Hu, Lawrence S. Moss, Ashish SabharwalAAAI 2020 · 被引用 152 次
- Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with TransformersColin Wei, Yining Chen, Tengyu MaNeurIPS 2022 · 被引用 117 次
相关 Paper
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
- Contextual Distortion Reveals Constituency: Masked Language Models are Implicit ParsersJiaxi Li, Wei LuACL 2023 · 被引用 3 次
- Unveiling the Black Box of PLMs with Semantic Anchors: Towards Interpretable Neural Semantic ParsingLunyiu Nie, Jiuding Sun, Yanlin Wang, Lun Du 等AAAI 2023 · 被引用 9 次
- Efficient Constituency Parsing by PointingThanh-Tung Nguyen, Xuan-Phi Nguyen, Shafiq R. Joty, Xiaoli LiACL 2020 · 被引用 11 次
- Evaluating the Impact of Model Scale for Compositional Generalization in Semantic ParsingLinlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi 等EMNLP 2022 · 被引用 21 次
