Language model acceptability judgements are not always robust to context
Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams
摘要
Targeted syntactic evaluations of language models ask whether models show stable preferences for syntactically acceptable content over minimal-pair unacceptable inputs. Our best syntactic evaluation datasets, however, provide substantially less linguistic context than models receive during pretraining. This mismatch raises an important question: how robust are models' syntactic judgements across different contexts? In this paper, we vary the input contexts based on: length, the types of syntactic phenomena it contains, and whether or not there are grammatical violations. We find that model judgements are generally robust when placed in randomly sampled linguistic contexts, but are unstable when contexts match the test stimuli in syntactic structure. Among all tested models (GPT-2 and five variants of OPT), we find that model performance is affected when we provided contexts with matching syntactic structure: performance significantly improves when contexts are acceptable, and it significantly declines when they are unacceptable. This effect is amplified by the length of the context, except for unrelated inputs. We show that these changes in model performance are not explainable by acceptability-preserving syntactic perturbations. This sensitivity to highly specific syntactic features of the context can only be explained by the models' implicit in-context learning abilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Evaluating Large Language Models on Controlled Generation TasksJiao Sun, Yufei Tian, Wangchunshu Zhou, Nan Xu 等EMNLP 2023 · 被引用 13 次
- Negative Pre-activations Differentiate SyntaxLinghao Kong, Angelina Ning, Micah Adler, Nir N ShavitICLR 2026 · 被引用 2 次
- Rapid Word Learning Through Meta In-Context LearningWentao Wang, Guangyuan Jiang, Tal Linzen, Brenden M. LakeEMNLP 2025
- Implicit Representations of Grammaticality in Language ModelsYingshan Susan Wang, Linlu Qiu, Zhaofeng Wu, Roger P. Levy 等ACL 2026
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel 等ACL 2022 · 被引用 1,494 次
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
- FewshotQA: A simple framework for few-shot learning of question answering tasks using pre-trained text-to-text modelsRakesh Chada, Pradeep NatarajanEMNLP 2021 · 被引用 36 次
相关 Paper
- Can In-context Learning Really Generalize to Out-of-distribution Tasks?Qixun Wang, Yifei Wang, Xianghua Ying, Yisen WangICLR 2025
- Parallel Structures in Pre-training Data Yield In-Context LearningYanda Chen, Chen Zhao, Zhou Yu, Kathleen R. McKeown 等ACL 2024 · 被引用 1 次
- How much pretraining data do language models need to learn syntax?Laura Pérez-Mayos, Miguel Ballesteros, Leo WannerEMNLP 2021 · 被引用 31 次
- On the Robustness of Language Encoders against Grammatical ErrorsFan Yin, Quanyu Long, Tao Meng, Kai-Wei ChangACL 2020 · 被引用 32 次
- A Targeted Assessment of Incremental Processing in Neural Language Models and HumansEthan Wilcox, Pranali Vani, Roger LevyACL 2021
