On the Robustness of Language Encoders against Grammatical Errors
Fan Yin, Quanyu Long, Tao Meng, Kai-Wei Chang
摘要
We conduct a thorough study to diagnose the behaviors of pre-trained language encoders (ELMo, BERT, and RoBERTa) when confronted with natural grammatical errors. Specifically, we collect real grammatical errors from non-native speakers and conduct adversarial attacks to simulate these errors on clean text data. We use this approach to facilitate debugging models on downstream applications. Results confirm that the performance of all tested models is affected but the degree of impact varies. To interpret model behaviors, we further design a linguistic acceptability task to reveal their abilities in identifying ungrammatical sentences and the position of errors. We find that fixed contextual encoders with a simple classifier trained on the prediction of sentence correctness are able to locate error positions. We also design a cloze test for BERT and discover that BERT captures the interaction between errors and specific tokens in context. Our results shed light on understanding the robustness and behaviors of language encoders against grammatical errors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- RuCoLA: Russian Corpus of Linguistic AcceptabilityVladislav Mikhailov, Tatiana Shamardina, Max Ryabinin, Alena Pestova 等EMNLP 2022 · 被引用 19 次
- ADDMU: Detection of Far-Boundary Adversarial Examples with Data and Model Uncertainty EstimationFan Yin, Yao Li, Cho-Jui Hsieh, Kai-Wei ChangEMNLP 2022 · 被引用 6 次
- Language Models Can be Efficiently Steered via Minimal Embedding Layer TransformationsDiogo Tavares, David Semedo, Alexander Rudnicky, João MagalhãesEMNLP 2025
相关 Paper
- Evaluating the Robustness of Neural Language Models to Input PerturbationsMilad Moradi, Matthias SamwaldEMNLP 2021 · 被引用 64 次
- Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding WordsJosef Klafka, Allyson EttingerACL 2020 · 被引用 2 次
- Positional Artefacts Propagate Through Masked Language Model EmbeddingsZiyang Luo, Artur Kulmizev, Xiaoxi MaoACL 2021
- BERT & Family Eat Word Salad: Experiments with Text UnderstandingAshim Gupta, Giorgi Kvernadze, Vivek SrikumarAAAI 2021 · 被引用 77 次
- Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQAIeva Staliunaite, Ignacio IacobacciEMNLP 2020 · 被引用 2 次
