Causal Inference over Visual-Semantic-Aligned Graph for Image Classification
Lei Meng, Xiangxian Li, Xiaoshuo Yan, Haokai Ma, Zhuang Qi, Wei Wu, Xiangxu Meng
摘要
Incorporating tagging information to regularize the representation learning of images usually leads to improved performance in image classification by aligning the visual features with the textual ones of higher discriminative power. Existing methods typically follow the predictive approach, which uses tags as the semantic labels for visual input to make predictions. However, they typically face the problem of handling the heterogeneity between modalities. In order to learn accurate visual-semantic mapping, this paper presents a visual-semantic causal association modeling framework termed VSCNet. It aligns visual regions with tags, uses a pre-learned hierarchy of visual and semantic exemplars to refine tag predictions and constructs an augmented heterogeneous graph to perform causal intervention. Specifically, the fine-grained visual-semantic alignment (FVA) module adaptively locates the semantic-intensive regions corresponding to tags. The heterogeneous association refinement (HAR) module associates the visual regions, semantic elements and pre-learned visual prototypes in a heterogeneous graph to filter the error predictions and enrich the information. The causal inference with graphical masking (CIM) module applies self-learned masks to discover the causal nodes and edges in the heterogeneous graph to address the spurious association, forming robust causal representations. Experimental results from two benchmarking datasets show that VSCNet effectively builds the visual-semantic associations from images and leads to better performance than the state-of-the-art methods with enriched predictive information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Class-wise Balancing Data Replay for Federated Class-Incremental LearningZhuang Qi, Ying-Peng Tang, Lei Meng, Han Yu 等NeurIPS 2025 · 被引用 11 次
- Global Prompt Refinement with Non-Interfering Attention Masking for One-Shot Federated LearningZhuang Qi, Pan Yu, Lei Meng, Sijin Zhou 等NeurIPS 2025 · 被引用 4 次
- Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal InterventionHaijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu 等NeurIPS 2025 · 被引用 2 次
- Prototype-based Causal Intervention for Multi-Label Image ClassificationYanmin Li, Zhilong Mao, Mao Wang, Lihua Liu 等CVPR 2026
- Explicit Modeling of Causal Factors and Confounders for Image ClassificationWei Wu, Lei Meng, Zhuang Qi, Zixuan Li 等AAAI 2026
它引用的顶会 Paper19
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- VanillaNet: the Power of Minimalism in Deep LearningHanting Chen, Yunhe Wang, Jianyuan Guo, Dacheng TaoNeurIPS 2023 · 被引用 228 次
- Causal Attention for Unbiased Visual RecognitionTan Wang, Chang Zhou, Qianru Sun, Hanwang ZhangICCV 2021 · 被引用 162 次
- Fine-Grained Semantically Aligned Vision-Language Pre-TrainingJuncheng Li, Xin He, Longhui Wei, Long Qian 等NeurIPS 2022 · 被引用 111 次
- Generalized Zero-Shot Learning via Disentangled RepresentationXiangyu Li, Zhe Xu, Kun Wei, Cheng DengAAAI 2021 · 被引用 88 次
相关 Paper
- Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text RetrievalYizhen Chen, Jie Wang, Lijian Lin, Zhongang Qi 等AAAI 2023 · 被引用 39 次
- Fine-grained Cross-modal Alignment Network for Text-Video RetrievalNing Han, Jingjing Chen, Guangyi Xiao, Hao Zhang 等ACM MM 2021 · 被引用 47 次
- Multi-View Differential Mixing and Graph-Guided Structural Region Selection for Cross-Modal AlignmentLinlin Ji, Li LiuAAAI 2026
- Learning Semantic-Specific Graph Representation for Multi-Label Image RecognitionTianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu 等ICCV 2019 · 被引用 347 次
- Tag2Text: Guiding Vision-Language Model via Image TaggingXinyu Huang, Youcai Zhang, Jinyu Ma, Weiwei Tian 等ICLR 2024 · 被引用 109 次
