Explore Spurious Correlations at the Concept Level in Language Models for Text Classification
Yuhang Zhou, Paiheng Xu, Xiaoyu Liu, Bang An, Wei Ai, Furong Huang
Abstract
Language models (LMs) have achieved notable success in numerous NLP tasks, employing both fine-tuning and in-context learning (ICL) methods. While language models demonstrate exceptional performance, they face robustness challenges due to spurious correlations arising from imbalanced label distributions in training data or ICL exemplars. Previous research has primarily concentrated on word, phrase, and syntax features, neglecting the concept level, often due to the absence of concept labels and difficulty in identifying conceptual content in input texts. This paper introduces two main contributions. First, we employ ChatGPT to assign concept labels to texts, assessing concept bias in models during fine-tuning or ICL on test data. We find that LMs, when encountering spurious correlations between a concept and a label in training or prompts, resort to shortcuts for predictions. Second, we introduce a data rebalancing technique that incorporates ChatGPTgenerated counterfactual data, thereby balancing label distribution and mitigating spurious correlations. Our method's efficacy, surpassing traditional token removal approaches, is validated through extensive testing. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd41dfa3-208e-4891-be32-718c2783859bCited by top-tier papers18
- The Best Instruction-Tuning Data are Those That FitDylan Zhang, Qirun Dai, Hao PengNeurIPS 2025 · 59 citations
- CSRec: Rethinking Sequential Recommendation from A Causal PerspectiveXiaoyu Liu, Jiaxin Yuan, Yuhang Zhou, Jingling Li et al.SIGIR 2025 · 7 citations
- Rectifying Shortcut Behaviors in Preference-based Reward LearningWenqian Ye, Guangtao Zheng, Aidong ZhangNeurIPS 2025 · 6 citations
- BOAD: Discovering Hierarchical Software Engineering Agents via Bandit OptimizationIris Xu, Guangtao Zeng, Zexue He, Charles Jin et al.ICLR 2026 · 5 citations
- Retrieving Counterfactuals Improves Visual In-Context LearningGuangzhi Xiong, Sanchit Sinha, Zhenghao He, Aidong ZhangCVPR 2026 · 3 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
Related papers
- ERICT: Enhancing Robustness by Identifying Concept Tokens in Zero-Shot Vision Language ModelsXinpeng Dong, Min Zhang, Didi Zhu, Ye Jun Jian et al.ICML 2025
- Unsupervised Concept Discovery Mitigates Spurious CorrelationsMd Rifat Arefin, Yan Zhang, Aristide Baratin, Francesco Locatello et al.ICML 2024 · 9 citations
- WRING Out The Bias: A Rotation-Based Alternative To Projection DebiasingWalter Gerych, Cassandra Parent, Quinn Perian, Rafiya Javed et al.ICLR 2026
- Interpret and Improve In-Context Learning via the Lens of Input-Label MappingsChenghao Sun, Zhen Huang, Yonggang Zhang, Le Lu et al.ACL 2025 · 1 citation
- Dually Self-Improved Counterfactual Data Augmentation Using Large Language ModelLuhao Zhang, Xinyu Zhang, Linmei Hu, Dandan Song et al.ACL 2025 · 1 citation
