Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
Chantal Shaib, Vinith M. Suriyakumar, Byron C. Wallace, Marzyeh Ghassemi
摘要
For an LLM to correctly respond to an instruction it must understand both the semantics and the domain (i.e., subject area) of a given task-instruction pair. However, syntax can also convey implicit information. Recent work shows that syntactic templates-frequent sequences of Part-of-Speech (PoS) tags-are prevalent in training data and often appear in model outputs. In this work we characterize syntactic templates, domain, and semantics in task-instruction pairs. We identify cases of spurious correlations between syntax and domain, where models learn to associate a domain with syntax during training; this can sometimes override prompt semantics. Using a synthetic training dataset, we find that the syntactic-domain correlation can lower performance (mean 0.51±0.06) on entity knowledge tasks in OLMo-2 models (1B-13B). We introduce an evaluation framework to detect this phenomenon in trained models, and show that it occurs on a subset of the FlanV2 dataset in open (OLMo-2-7B; Llama-4-Maverick), and closed (GPT-4o) models. Finally, we present a case study on the implications for LLM security, showing that unintended syntactic-domain correlations can be used to bypass refusals in OLMo-2-7B Instruct and GPT-4o. Our findings highlight two needs: (1) to explicitly test for syntactic-domain correlations, and (2) to ensure syntactic diversity in training data, specifically within domains, to prevent such spurious correlations.
Content Warning: This paper contains examples of harmful language.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Differential syntactic and semantic encoding in LLMsSantiago Acevedo, Alessandro Laio, Marco BaroniICML 2026 · 被引用 7 次
- Causal Fine-Tuning under Latent Confounded ShiftJialin Yu, Yuxiang Zhou, Haoxuan Li, Junchi Yu 等ICML 2026
它引用的顶会 Paper15
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 被引用 454 次
- Learning De-biased Representations with Biased RepresentationsHyojin Bahng, Sanghyuk Chun, Sangdoo Yun, Jaegul Choo 等ICML 2020 · 被引用 332 次
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger 等NeurIPS 2024 · 被引用 247 次
相关 Paper
- Detection and Measurement of Syntactic Templates in Generated TextChantal Shaib, Yanai Elazar, Junyi Jessy Li, Byron C. WallaceEMNLP 2024 · 被引用 3 次
- StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt GenerationHuawei Zheng, Xinqi Jiang, Sen Yang, Shouling Ji 等ACL 2026 · 被引用 1 次
- Defenses Against Prompt Attacks Learn Surface HeuristicsShawn Li, Chenxiao Yu, Zhiyu Ni, Hao Li 等ACL 2026 · 被引用 8 次
- The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMsSergey Berezin, Reza Farahbakhsh, Noël CrespiACL 2025
- Underspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun ResolutionEmily McMilinAAAI 2024 · 被引用 1 次
