Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models
Chantal Shaib, Vinith M. Suriyakumar, Byron C. Wallace, Marzyeh Ghassemi
Abstract
For an LLM to correctly respond to an instruction it must understand both the semantics and the domain (i.e., subject area) of a given task-instruction pair. However, syntax can also convey implicit information. Recent work shows that syntactic templates-frequent sequences of Part-of-Speech (PoS) tags-are prevalent in training data and often appear in model outputs. In this work we characterize syntactic templates, domain, and semantics in task-instruction pairs. We identify cases of spurious correlations between syntax and domain, where models learn to associate a domain with syntax during training; this can sometimes override prompt semantics. Using a synthetic training dataset, we find that the syntactic-domain correlation can lower performance (mean 0.51±0.06) on entity knowledge tasks in OLMo-2 models (1B-13B). We introduce an evaluation framework to detect this phenomenon in trained models, and show that it occurs on a subset of the FlanV2 dataset in open (OLMo-2-7B; Llama-4-Maverick), and closed (GPT-4o) models. Finally, we present a case study on the implications for LLM security, showing that unintended syntactic-domain correlations can be used to bypass refusals in OLMo-2-7B Instruct and GPT-4o. Our findings highlight two needs: (1) to explicitly test for syntactic-domain correlations, and (2) to ensure syntactic diversity in training data, specifically within domains, to prevent such spurious correlations.
Content Warning: This paper contains examples of harmful language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be560376-d0dc-4bf9-bc87-6535ae0dc806Cited by top-tier papers2
- Differential syntactic and semantic encoding in LLMsSantiago Acevedo, Alessandro Laio, Marco BaroniICML 2026 · 7 citations
- Causal Fine-Tuning under Latent Confounded ShiftJialin Yu, Yuxiang Zhou, Haoxuan Li, Junchi Yu et al.ICML 2026
Builds on15
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 454 citations
- Learning De-biased Representations with Biased RepresentationsHyojin Bahng, Sanghyuk Chun, Sangdoo Yun, Jaegul Choo et al.ICML 2020 · 332 citations
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger et al.NeurIPS 2024 · 247 citations
Related papers
- Detection and Measurement of Syntactic Templates in Generated TextChantal Shaib, Yanai Elazar, Junyi Jessy Li, Byron C. WallaceEMNLP 2024 · 3 citations
- StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt GenerationHuawei Zheng, Xinqi Jiang, Sen Yang, Shouling Ji et al.ACL 2026 · 1 citation
- Defenses Against Prompt Attacks Learn Surface HeuristicsShawn Li, Chenxiao Yu, Zhiyu Ni, Hao Li et al.ACL 2026 · 8 citations
- The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMsSergey Berezin, Reza Farahbakhsh, Noël CrespiACL 2025
- Underspecification in Language Modeling Tasks: A Causality-Informed Study of Gendered Pronoun ResolutionEmily McMilinAAAI 2024 · 1 citation
