Synthesizing Counterfactual Samples for Effective Image-Text Matching
Hao Wei, Shuhui Wang, Xinzhe Han, Zhe Xue, Bin Ma, Xiaoming Wei, Xiaolin Wei
Abstract
Image-text matching is a fundamental research topic bridging vision and language. Recent works use hard negative mining to capture the multiple correspondences between visual and textual domains. Unfortunately, the truly informative negative samples are quite sparse in the training data, which are hard to obtain only in a randomly sampled mini-batch. Motivated by causal inference, we aim to overcome this shortcoming by carefully analyzing the analogy between hard negative mining and causal effects optimizing. Further, we propose Counterfactual Matching (CFM) framework for more effective image-text correspondence mining. CFM contains three major components, , Gradient-Guided Feature Selection for automatic casual factor identification, Self-Exploration for causal factor completeness, and Self-Adjustment for counterfactual sample synthesis. Compared with traditional hard negative mining, our method largely alleviates the over-fitting phenomenon and effectively captures the fine-grained correlations between image and text modality. We evaluate our CFM in combination with three state-of-the-art image-text matching architectures. Quantitative and qualitative experiments conducted on two publicly available datasets demonstrate its strong generality and effectiveness. Code is available at: https://github.com/weihao20/cfm.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- Identification of Necessary Semantic Undertakers in the Causal View for Image-Text MatchingHuatian Zhang, Lei Zhang, Kun Zhang, Zhendong MaoAAAI 2024 · 12 citations
- Towards Deconfounded Image-Text Matching with Causal InferenceWenhui Li, Xinqi Su, Dan Song, Lanjun Wang et al.ACM MM 2023 · 12 citations
- Expanding the Scope of Negatives: Boosting Image-Text Matching with Negatives Distribution Guided LearningZhao Zhou, Weizhong Zhang, Xiangcheng Du, Yingbin Zheng et al.AAAI 2025
- Camouflage-aware Image-Text Retrieval via Expert CollaborationYao Jiang, Zhongkuan Mao, Xuan Wu, Keren Fu et al.CVPR 2026
Related papers
- CF-VLM: CounterFactual Vision-Language Fine-tuningJusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang et al.NeurIPS 2025 · 71 citations
- When Hard Negatives Hurt: Bridging the Generative Discriminative Gap in Hard Negative Synthesis for RetrievalZhicheng Zhang, Jiwei Tang, Kuicai Dong, Xiaopeng Li et al.KDD 2026 · 1 citation
- Semantic-Aware Hard Negative Mining for Medical Vision-Language Contrastive PretrainingYongxin Li, Ying Cheng, Yaning Pan, Wen He et al.ACM MM 2025 · 2 citations
- Counterfactual Contrastive Learning for Weakly-Supervised Vision-Language GroundingZhu Zhang, Zhou Zhao, Zhijie Lin, Jieming Zhu et al.NeurIPS 2020 · 74 citations
- Uncovering Main Causalities for Long-tailed Information ExtractionGuoshun Nan, Jiaqi Zeng, Rui Qiao, Zhijiang Guo et al.EMNLP 2021 · 39 citations
