Causal Inference over Visual-Semantic-Aligned Graph for Image Classification
Lei Meng, Xiangxian Li, Xiaoshuo Yan, Haokai Ma, Zhuang Qi, Wei Wu, Xiangxu Meng
Abstract
Incorporating tagging information to regularize the representation learning of images usually leads to improved performance in image classification by aligning the visual features with the textual ones of higher discriminative power. Existing methods typically follow the predictive approach, which uses tags as the semantic labels for visual input to make predictions. However, they typically face the problem of handling the heterogeneity between modalities. In order to learn accurate visual-semantic mapping, this paper presents a visual-semantic causal association modeling framework termed VSCNet. It aligns visual regions with tags, uses a pre-learned hierarchy of visual and semantic exemplars to refine tag predictions and constructs an augmented heterogeneous graph to perform causal intervention. Specifically, the fine-grained visual-semantic alignment (FVA) module adaptively locates the semantic-intensive regions corresponding to tags. The heterogeneous association refinement (HAR) module associates the visual regions, semantic elements and pre-learned visual prototypes in a heterogeneous graph to filter the error predictions and enrich the information. The causal inference with graphical masking (CIM) module applies self-learned masks to discover the causal nodes and edges in the heterogeneous graph to address the spurious association, forming robust causal representations. Experimental results from two benchmarking datasets show that VSCNet effectively builds the visual-semantic associations from images and leads to better performance than the state-of-the-art methods with enriched predictive information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Class-wise Balancing Data Replay for Federated Class-Incremental LearningZhuang Qi, Ying-Peng Tang, Lei Meng, Han Yu et al.NeurIPS 2025 · 11 citations
- Global Prompt Refinement with Non-Interfering Attention Masking for One-Shot Federated LearningZhuang Qi, Pan Yu, Lei Meng, Sijin Zhou et al.NeurIPS 2025 · 4 citations
- Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal InterventionHaijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu et al.NeurIPS 2025 · 2 citations
- Prototype-based Causal Intervention for Multi-Label Image ClassificationYanmin Li, Zhilong Mao, Mao Wang, Lihua Liu et al.CVPR 2026
- Explicit Modeling of Causal Factors and Confounders for Image ClassificationWei Wu, Lei Meng, Zhuang Qi, Zixuan Li et al.AAAI 2026
Builds on19
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- VanillaNet: the Power of Minimalism in Deep LearningHanting Chen, Yunhe Wang, Jianyuan Guo, Dacheng TaoNeurIPS 2023 · 228 citations
- Causal Attention for Unbiased Visual RecognitionTan Wang, Chang Zhou, Qianru Sun, Hanwang ZhangICCV 2021 · 162 citations
- Fine-Grained Semantically Aligned Vision-Language Pre-TrainingJuncheng Li, Xin He, Longhui Wei, Long Qian et al.NeurIPS 2022 · 111 citations
- Generalized Zero-Shot Learning via Disentangled RepresentationXiangyu Li, Zhe Xu, Kun Wei, Cheng DengAAAI 2021 · 88 citations
Related papers
- Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text RetrievalYizhen Chen, Jie Wang, Lijian Lin, Zhongang Qi et al.AAAI 2023 · 39 citations
- Fine-grained Cross-modal Alignment Network for Text-Video RetrievalNing Han, Jingjing Chen, Guangyi Xiao, Hao Zhang et al.ACM MM 2021 · 47 citations
- Multi-View Differential Mixing and Graph-Guided Structural Region Selection for Cross-Modal AlignmentLinlin Ji, Li LiuAAAI 2026
- Learning Semantic-Specific Graph Representation for Multi-Label Image RecognitionTianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu et al.ICCV 2019 · 347 citations
- Tag2Text: Guiding Vision-Language Model via Image TaggingXinyu Huang, Youcai Zhang, Jinyu Ma, Weiwei Tian et al.ICLR 2024 · 109 citations
