On the Robustness of Language Encoders against Grammatical Errors
Fan Yin, Quanyu Long, Tao Meng, Kai-Wei Chang
Abstract
We conduct a thorough study to diagnose the behaviors of pre-trained language encoders (ELMo, BERT, and RoBERTa) when confronted with natural grammatical errors. Specifically, we collect real grammatical errors from non-native speakers and conduct adversarial attacks to simulate these errors on clean text data. We use this approach to facilitate debugging models on downstream applications. Results confirm that the performance of all tested models is affected but the degree of impact varies. To interpret model behaviors, we further design a linguistic acceptability task to reveal their abilities in identifying ungrammatical sentences and the position of errors. We find that fixed contextual encoders with a simple classifier trained on the prediction of sentence correctness are able to locate error positions. We also design a cloze test for BERT and discover that BERT captures the interaction between errors and specific tokens in context. Our results shed light on understanding the robustness and behaviors of language encoders against grammatical errors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5296f4e6-264a-4cee-909a-d8778329bf8eCited by top-tier papers3
- RuCoLA: Russian Corpus of Linguistic AcceptabilityVladislav Mikhailov, Tatiana Shamardina, Max Ryabinin, Alena Pestova et al.EMNLP 2022 · 19 citations
- ADDMU: Detection of Far-Boundary Adversarial Examples with Data and Model Uncertainty EstimationFan Yin, Yao Li, Cho-Jui Hsieh, Kai-Wei ChangEMNLP 2022 · 6 citations
- Language Models Can be Efficiently Steered via Minimal Embedding Layer TransformationsDiogo Tavares, David Semedo, Alexander Rudnicky, João MagalhãesEMNLP 2025
Related papers
- Evaluating the Robustness of Neural Language Models to Input PerturbationsMilad Moradi, Matthias SamwaldEMNLP 2021 · 64 citations
- Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding WordsJosef Klafka, Allyson EttingerACL 2020 · 2 citations
- Positional Artefacts Propagate Through Masked Language Model EmbeddingsZiyang Luo, Artur Kulmizev, Xiaoxi MaoACL 2021
- BERT & Family Eat Word Salad: Experiments with Text UnderstandingAshim Gupta, Giorgi Kvernadze, Vivek SrikumarAAAI 2021 · 77 citations
- Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQAIeva Staliunaite, Ignacio IacobacciEMNLP 2020 · 2 citations
