How is BERT surprised? Layerwise detection of linguistic anomalies
Bai Li, Zining Zhu, Guillaume Thomas, Yang Xu, Frank Rudzicz
摘要
Transformer language models have shown remarkable ability in detecting when a word is anomalous in context, but likelihood scores offer no information about the cause of the anomaly. In this work, we use Gaussian models for density estimation at intermediate layers of three language models (BERT, RoBERTa, and XLNet), and evaluate our method on BLiMP, a grammaticality judgement benchmark. In lower layers, surprisal is highly correlated to low token frequency, but this correlation diminishes in upper layers. Next, we gather datasets of morphosyntactic, semantic, and commonsense anomalies from psycholinguistic studies; we find that the best performing model RoBERTa exhibits surprisal in earlier layers when the anomaly is morphosyntactic than when it is semantic, while commonsense anomalies do not exhibit surprisal at any intermediate layer. These results suggest that language models employ separate mechanisms to detect different types of linguistic anomalies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Closer Look at How Fine-tuning Changes BERTYichu Zhou, Vivek SrikumarACL 2022 · 被引用 84 次
- Neural reality of argument structure constructionsBai Li, Zining Zhu, Guillaume Thomas, Frank Rudzicz 等ACL 2022 · 被引用 38 次
- Predicting Fine-Tuning Performance with ProbingZining Zhu, Soroosh Shahtalebi, Frank RudziczEMNLP 2022 · 被引用 6 次
它引用的顶会 Paper7
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 被引用 167 次
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
- Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain DetectionAlexander Podolskiy, Dmitry Lipin, Andrey Bout, Ekaterina Artemova 等AAAI 2021 · 被引用 100 次
- Investigating Word-Class Distributions in Word Vector SpacesRyohei Sasano, Anna KorhonenACL 2020 · 被引用 7 次
- Intrinsic Probing through Dimension SelectionLucas Torroba Hennigen, Adina Williams, Ryan CotterellEMNLP 2020 · 被引用 3 次
相关 Paper
- Probing Pretrained Language Models for Lexical SemanticsIvan Vulic, Edoardo Maria Ponti, Robert Litschko, Goran Glavas 等EMNLP 2020 · 被引用 26 次
- When Do You Need Billions of Words of Pretraining Data?Yian Zhang, Alex Warstadt, Xiaocheng Li, Samuel R. BowmanACL 2021
- Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQAIeva Staliunaite, Ignacio IacobacciEMNLP 2020 · 被引用 2 次
- On the Robustness of Language Encoders against Grammatical ErrorsFan Yin, Quanyu Long, Tao Meng, Kai-Wei ChangACL 2020 · 被引用 32 次
- The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative CorrelativeLeonie Weissweiler, Valentin Hofmann, Abdullatif Köksal, Hinrich SchützeEMNLP 2022 · 被引用 13 次
