Learning Stable Classifiers by Transferring Unstable Features
Yujia Bao, Shiyu Chang, Regina Barzilay
Abstract
While unbiased machine learning models are essential for many applications, bias is a human-defined concept that can vary across tasks. Given only input-label pairs, algorithms may lack sufficient information to distinguish stable (causal) features from unstable (spurious) features. However, related tasks often share similar biases -- an observation we may leverage to develop stable classifiers in the transfer setting. In this work, we explicitly inform the target classifier about unstable features in the source tasks. Specifically, we derive a representation that encodes the unstable features by contrasting different data environments in the source task. We achieve robustness by clustering data of the target task according to this representation and minimizing the worst-case risk across these clusters. We evaluate our method on both text and image classifications. Empirical results demonstrate that our algorithm is able to maintain robustness on the target task for both synthetically generated environments and real-world environments. Our code is available at https://github.com/YujiaBao/Tofu.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7aa0b6d3-0131-4427-8148-8dc4941f82d3Cited by top-tier papers2
- SelecMix: Debiased Learning by Contradicting-pair SamplingInwoo Hwang, Sangjun Lee, Yunhyeok Kwak, Seong Joon Oh et al.NeurIPS 2022 · 43 citations
- A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies OthersZhiheng Li, Ivan Evtimov, Albert Gordo, Caner Hazirbas et al.CVPR 2023
Builds on20
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 454 citations
- Domain Generalization Using a Mixture of Multiple Latent DomainsToshihiko Matsuura, Tatsuya HaradaAAAI 2020 · 355 citations
- No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification ProblemsNimit Sharad Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu et al.NeurIPS 2020 · 316 citations
Related papers
- Causal Transportability for Visual RecognitionChengzhi Mao, Kevin Xia, James Wang, Hao Wang et al.CVPR 2022 · 27 citations
- TACIT: A Target-Agnostic Feature Disentanglement Framework for Cross-Domain Text ClassificationRui Song, Fausto Giunchiglia, Yingji Li, Mingjie Tian et al.AAAI 2024 · 10 citations
- Examining and Combating Spurious Features under Distribution ShiftChunting Zhou, Xuezhe Ma, Paul Michel, Graham NeubigICML 2021 · 78 citations
- Adv-SSL: Adversarial Self-Supervised Representation Learning with Theoretical GuaranteesChenguang Duan, Yuling Jiao, Huazhen Lin, Wensen Ma et al.NeurIPS 2025 · 1 citation
- Spuriosity Didn't Kill the Classifier: Using Invariant Predictions to Harness Spurious FeaturesCian Eastwood, Shashank Singh, Andrei Liviu Nicolicioiu, Marin Vlastelica Pogancic et al.NeurIPS 2023 · 29 citations
