Probing as Quantifying Inductive Bias
Alexander Immer, Lucas Torroba Hennigen, Vincent Fortuin, Ryan Cotterell
摘要
Pre-trained contextual representations have led to dramatic performance improvements on a range of downstream tasks. Such performance improvements have motivated researchers to quantify and understand the linguistic information encoded in these representations. In general, researchers quantify the amount of linguistic information through probing, an endeavor which consists of training a supervised model to predict a linguistic property directly from the contextual representations. Unfortunately, this definition of probing has been subject to extensive criticism in the literature, and has been observed to lead to paradoxical and counter-intuitive results. In the theoretical portion of this paper, we take the position that the goal of probing ought to be measuring the amount of inductive bias that the representations encode on a specific task. We further describe a Bayesian framework that operationalizes this goal and allows us to quantify the representations' inductive bias. In the empirical portion of the paper, we apply our framework to a variety of NLP tasks. Our results suggest that our proposed framework alleviates many previous problems found in probing. Moreover, we are able to offer concrete evidence that-for some tasks-fastText can offer a better inductive bias than BERT. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Invariance Learning in Deep Neural Networks with Differentiable Laplace ApproximationsAlexander Immer, Tycho F. A. van der Ouderaa, Gunnar Rätsch, Vincent Fortuin 等NeurIPS 2022 · 被引用 56 次
- Stochastic Marginal Likelihood Gradients using Neural Tangent KernelsAlexander Immer, Tycho F. A. van der Ouderaa, Mark van der Wilk, Gunnar Rätsch 等ICML 2023 · 被引用 17 次
- Improving Neural Additive Models with Bayesian PrinciplesKouroche Bouchiat, Alexander Immer, Hugo Yèche, Gunnar Rätsch 等ICML 2024 · 被引用 17 次
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang 等ACL 2026 · 被引用 5 次
- The Distributional Hypothesis Does Not Fully Explain the Benefits of Masked Language Model PretrainingTing-Rui Chiang, Dani YogatamaEMNLP 2023 · 被引用 1 次
它引用的顶会 Paper11
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
- Repulsive Deep Ensembles are BayesianFrancesco D'Angelo, Vincent FortuinNeurIPS 2021 · 被引用 141 次
- Predicting Inductive Biases of Pre-Trained ModelsCharles Lovering, Rohan Jha, Tal Linzen, Ellie PavlickICLR 2021 · 被引用 70 次
- On the Pitfalls of Analyzing Individual Neurons in Language ModelsOmer Antverg, Yonatan BelinkovICLR 2022 · 被引用 64 次
相关 Paper
- Intrinsic Probing through Dimension SelectionLucas Torroba Hennigen, Adina Williams, Ryan CotterellEMNLP 2020 · 被引用 3 次
- A Latent-Variable Model for Intrinsic ProbingKarolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell 等AAAI 2023 · 被引用 6 次
- Information-Theoretic Probing for Linguistic StructureTiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod 等ACL 2020 · 被引用 21 次
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 被引用 158 次
- A matter of framing: The impact of linguistic formalism on probing resultsIlia Kuznetsov, Iryna GurevychEMNLP 2020 · 被引用 1 次
