Explicit Modeling of Causal Factors and Confounders for Image Classification
Wei Wu, Lei Meng, Zhuang Qi, Zixuan Li, Yachong Zhang, Xiaoshuo Yan, Xiangxu Meng
Abstract
Causal inference has emerged as a promising approach for identifying decisive semantic factors and eliminating spurious correlations in visual representation learning. However, most existing methods rely on latent, data-driven confounder modeling, normally attributing the source of bias to background information while neglecting object-level semantic confusions that commonly occur in complex scenes. This limits their effectiveness in disentangling causal factors from confounding semantics. To address this challenge, we propose an explicit modeling approach for both causal factors and confounders, termed Explicit Modeling Causal Model (EMCM). The proposed framework consists of three key components. The Features Stability Estimation module explicitly models the relationship between visual semantics and class labels by leveraging clustering patterns to perform class-aware separation of causal and confounding factors. It produces class-specific causal factors and confounding factors linked to ambiguous categories. Subsequently, the Discriminative Features Enhancing module integrates causal factors into fused patch features via front-door intervention for stable semantics. In parallel, the Explicit Confounder Modeling and Debiasing Module learns confounders under clear label guidance and derives debiased context features by TDE modeling. This framework leverages two complementary causal perspectives to construct a unified semantic representation that facilitates improved generalization. Extensive experiments on two datasets demonstrate that EMCM effectively disentangles causal and confounding factors in complex scenarios, consistently outperforming state-of-the-art causal debiasing methods and text-guided methods in all metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e00f7a7-edd8-4ab2-8b2a-6e5763bfdd49Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Vision-Language Pre-Training with Triple Contrastive LearningJinyu Yang, Jiali Duan, Son Tran, Yi Xu et al.CVPR 2022 · 266 citations
- SoftCLIP: Softer Cross-Modal Alignment Makes CLIP StrongerYuting Gao, Jinfeng Liu, Zihan Xu, Tong Wu et al.AAAI 2024 · 80 citations
- Robust Cross-Modal Representation Learning with Progressive Self-DistillationAlex Andonian, Shixing Chen, Raffay HamidCVPR 2022 · 43 citations
- Contextual Debiasing for Visual Recognition with Causal MechanismsRuyang Liu, Hao Liu, Ge Li, Haodi Hou et al.CVPR 2022 · 42 citations
Related papers
- C-Disentanglement: Discovering Causally-Independent Generative Factors under an Inductive Bias of ConfounderXiaoyu Liu, Jiaxin Yuan, Bang An, Yuancheng Xu et al.NeurIPS 2023 · 13 citations
- Show, Deconfound and Tell: Image Captioning with Causal InferenceBing Liu, Dong Wang, Xu Yang, Yong Zhou et al.CVPR 2022 · 66 citations
- Visual Representation Learning through Causal Intervention for Controllable Image EditingShanshan Huang, Haoxuan Li, Chunyuan Zheng, Lei Wang et al.CVPR 2025
- Deconfounded Multimodal Learning for Spatio-temporal Video GroundingJiawei Wang, Zhanchang Ma, Da Cao, Yuquan Le et al.ACM MM 2023 · 7 citations
- Modeling Event-level Causal Representation for Video ClassificationYuqing Wang, Lei Meng, Haokai Ma, Yuqing Wang et al.ACM MM 2024 · 3 citations
