DUET: Cross-Modal Semantic Grounding for Contrastive Zero-Shot Learning
Zhuo Chen, Yufeng Huang, Jiaoyan Chen, Yuxia Geng, Wen Zhang, Yin Fang, Jeff Z. Pan, Huajun Chen
摘要
Zero-shot learning (ZSL) aims to predict unseen classes whose samples have never appeared during training. As annotations for class-level visual characteristics, attributes are among the most effective and widely used semantic information for zero-shot image classification. However, the current methods often fail to discriminate those subtle visual distinctions between images due to not only the lack of fine-grained annotations, but also the issues of attribute imbalance and co-occurrence. In this paper, we present a transformer-based end-to-end ZSL method named DUET, which integrates latent semantic knowledge from the pretrained language models (PLMs) via a self-supervised multimodal learning paradigm. Specifically, we (1) developed a cross-modal semantic grounding network to investigate the model's capability of disentangling semantic attributes from the images; (2) applied an attribute-level contrastive learning strategy to further enhance the model's discrimination on fine-grained visual characteristics against the attribute cooccurrence and imbalance; (3) proposed a multi-task learning policy for considering multi-model objectives. We find that DUET can achieve state-of-the-art performance on three standard ZSL benchmarks and a knowledge graph equipped ZSL benchmark, and that its components are effective and its predictions are interpretable.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-Modal Structured RepresentationsYufeng Huang, Jiji Tang, Zhuo Chen, Rongsheng Zhang 等AAAI 2024 · 被引用 65 次
- Data Distribution Distilled Generative Model for Generalized Zero-Shot RecognitionYijie Wang, Mingjian Hong, Luwen Huangfu, Sheng HuangAAAI 2024 · 被引用 21 次
- CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot LearningHaojian Huang, Xiaozhen Qiao, Zhuo Chen, Haodong Chen 等ACM MM 2024 · 被引用 12 次
- SVIP: Semantically Contextualized Visual Patches for Zero-Shot LearningZhi Chen, Zecheng Zhao, Jingcai Guo, Jingjing Li 等ICCV 2025 · 被引用 8 次
- Enhancing Visual Continual Learning with Language-Guided SupervisionBolin Ni, Hongbo Zhao, Chenghao Zhang, Ke Hu 等CVPR 2024 · 被引用 8 次
它引用的顶会 Paper20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
相关 Paper
- Progressive Semantic-Guided Vision Transformer for Zero-Shot LearningShiming Chen, Wenjin Hou, Salman H. Khan, Fahad Shahbaz KhanCVPR 2024
- TransZero: Attribute-Guided Transformer for Zero-Shot LearningShiming Chen, Ziming Hong, Yang Liu, Guo-Sen Xie 等AAAI 2022 · 被引用 185 次
- Recognizing Unseen Objects via Multimodal Intensive Knowledge Graph PropagationLikang Wu, Zhi Li, Hongke Zhao, Zhefeng Wang 等KDD 2023 · 被引用 4 次
- I2DFormer: Learning Image to Document Attention for Zero-Shot Image ClassificationMuhammad Ferjad Naeem, Yongqin Xian, Luc Van Gool, Federico TombariNeurIPS 2022 · 被引用 63 次
- Attend and Enrich: Enhanced Visual Prompt for Zero-Shot LearningMan Liu, Huihui Bai, Feng Li, Chunjie Zhang 等AAAI 2025 · 被引用 3 次
