CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
Kyohoon Jin, Juhwan Choi, Jungmin Yun, Junho Lee, Soojin Jang, YoungBin Kim
Abstract
Deep learning models often learn and exploit spurious correlations in training data, using these non-target features to inform their predictions. Such reliance leads to performance degradation and poor generalization on unseen data. To address these limitations, we introduce a more general form of counterfactual data augmentation, termed counterbias data augmentation, which simultaneously tackles multiple biases (e.g., gender bias, simplicity bias) and enhances out-of-distribution robustness. We present CoBA: CounterBias Augmentation, a unified framework that operates at the semantic triple level: first decomposing text into subject-predicate-object triples, then selectively modifying these triples to disrupt spurious correlations. By reconstructing the text from these adjusted triples, CoBA generates counterbias data that mitigates spurious patterns. Through extensive experiments, we demonstrate that CoBA not only improves downstream task performance, but also effectively reduces biases and strengthens out-of-distribution resilience, offering a versatile and robust solution to the challenges posed by spurious correlations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66d4c9cd-4de7-4953-93dc-fbb79122fcd4Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- ReAct: Out-of-distribution Detection With Rectified ActivationsYiyou Sun, Chuan Guo, Yixuan LiNeurIPS 2021 · 733 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
Related papers
- C2L: Causally Contrastive Learning for Robust Text ClassificationSeungtaek Choi, Myeongho Jeong, Hojae Han, Seung-won HwangAAAI 2022 · 52 citations
- Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity TypingNan Xu, Fei Wang, Bangzheng Li, Mingtao Dong et al.EMNLP 2022 · 6 citations
- BiasAdv: Bias-Adversarial Augmentation for Model DebiasingJongin Lim, Youngdong Kim, Byungjai Kim, Chanho Ahn et al.CVPR 2023
- Causal-structure Driven Augmentations for Text OOD GeneralizationAmir Feder, Yoav Wald, Claudia Shi, Suchi Saria et al.NeurIPS 2023 · 10 citations
- Counterfactual Inference for Text Classification DebiasingChen Qian, Fuli Feng, Lijie Wen, Chunping Ma et al.ACL 2021
