Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERT
Zhiyong Wu, Yun Chen, Ben Kao, Qun Liu
Abstract
By introducing a small set of additional parameters, a probe learns to solve specific linguistic tasks (e.g., dependency parsing) in a supervised manner using feature representations (e.g., contextualized embeddings). The effectiveness of such probing tasks is taken as evidence that the pre-trained model encodes linguistic knowledge. However, this approach of evaluating a language model is undermined by the uncertainty of the amount of knowledge that is learned by the probe itself. Complementary to those works, we propose a parameter-free probing technique for analyzing pre-trained language models (e.g., BERT). Our method does not require direct supervision from the probing tasks, nor do we introduce additional parameters to the probing process. Our experiments on BERT show that syntactic trees recovered from BERT using our method are significantly better than linguistically-uninformed baselines. We further feed the empirically induced dependency structures into a downstream sentiment classification task and find its improvement compatible with or even superior to a human-designed dependency schema. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers32
- The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer EncodersHan He, Jinho D. ChoiEMNLP 2021 · 111 citations
- A Closer Look at How Fine-tuning Changes BERTYichu Zhou, Vivek SrikumarACL 2022 · 84 citations
- OA-Mine: Open-World Attribute Mining for E-Commerce Products with Weak SupervisionXinyang Zhang, Chenwei Zhang, Xian Li, Xin Luna Dong et al.WWW 2022 · 36 citations
- Analyzing How BERT Performs Entity MatchingMatteo Paganelli, Francesco Del Buono, Andrea Baraldi, Francesco GuerraVLDB 2022 · 35 citations
- Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammarsKaiyue Wen, Yuchen Li, Bingbin Liu, Andrej RisteskiNeurIPS 2023 · 32 citations
Related papers
- Probing as Quantifying Inductive BiasAlexander Immer, Lucas Torroba Hennigen, Vincent Fortuin, Ryan CotterellACL 2022
- A Latent-Variable Model for Intrinsic ProbingKarolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell et al.AAAI 2023 · 6 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Probing for Labeled Dependency TreesMax Müller-Eberstein, Rob van der Goot, Barbara PlankACL 2022 · 10 citations
- Exploring the Role of BERT Token Representations to Explain Sentence Probing ResultsHosein Mohebbi, Ali Modarressi, Mohammad Taher PilehvarEMNLP 2021 · 14 citations
