Kernel-Whitening: Overcome Dataset Bias with Isotropic Sentence Embedding
Songyang Gao, Shihan Dou, Qi Zhang, Xuanjing Huang
摘要
Dataset bias has attracted increasing attention recently for its detrimental effect on the generalization ability of fine-tuned models. The current mainstream solution is designing an additional shallow model to pre-identify biased instances. However, such two-stage methods scale up the computational complexity of training process and obstruct valid feature information while mitigating bias. To address this issue, we utilize the representation normalization method which aims at disentangling the correlations between features of encoded sentences. We find it also promising in eliminating the bias problem by providing isotropic data distribution. We further propose Kernel-Whitening, a Nyström kernel approximation method to achieve more thorough debiasing on nonlinear spurious correlations. Our framework is end-to-end with similar time consumption to fine-tuning. Experiments show that Kernel-Whitening significantly improves the performance of BERT on out-of-distribution datasets while maintaining in-distribution accuracy. 1 A positive definite symmetric matrix.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FairFlow: Mitigating Dataset Biases through Undecided Learning for Natural Language UnderstandingJiali Cheng, Hadi AmiriEMNLP 2024 · 被引用 2 次
- REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuningSeungmin Lee, Jeonghwan Lee, Hyunkuk Lim, Sejoon Kim 等ACL 2026
它引用的顶会 Paper6
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Learning from others' mistakes: Avoiding dataset biases without modeling themVictor Sanh, Thomas Wolf, Yonatan Belinkov, Alexander M. RushICLR 2021 · 被引用 123 次
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu 等NeurIPS 2021 · 被引用 102 次
- Uncertainty Calibration for Ensemble-Based Debiasing MethodsRuibin Xiong, Yimeng Chen, Liang Pang, Xueqi Cheng 等NeurIPS 2021 · 被引用 23 次
- Towards Debiasing NLU Models from Unknown BiasesPrasetya Ajie Utama, Nafise Sadat Moosavi, Iryna GurevychEMNLP 2020 · 被引用 3 次
相关 Paper
- IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationWenxuan Zhou, Bill Yuchen Lin, Xiang RenAAAI 2021 · 被引用 29 次
- FairFil: Contrastive Neural Debiasing Method for Pretrained Text EncodersPengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si 等ICLR 2021 · 被引用 50 次
- Debiasing Pretrained Text Encoders by Paying Attention to Paying AttentionYacine Gaci, Boualem Benatallah, Fabio Casati, Khalid BenabdeslemEMNLP 2022 · 被引用 12 次
- Towards Debiasing Sentence RepresentationsPaul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim 等ACL 2020 · 被引用 149 次
- Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of SamplesWeiwei Li, Junzhuo Liu, Yuanyuan Ren, Yuchen Zheng 等CVPR 2025
