M3R: Masked Token Mixup and Cross-Modal Reconstruction for Zero-Shot Learning
Peng Zhao, Qiangchang Wang, Yilong Yin
摘要
In the zero-shot learning (ZSL), learned representation spaces are often biased toward seen classes, thus limiting the ability to predict previously unseen classes. In this paper, we propose Masked token Mixup and cross-Modal Reconstruction for zero-shot learning, termed as M3R, which can significantly alleviate the bias toward seen classes. The M3R mainly consists of Random Token Mixup (RTM), Unseen Class Detection (UCD), and Hard Cross-modal Reconstruction (HCR). Firstly, mappings without proper adaptations to unseen classes would cause the bias toward seen classes. To address this issue, the RTM is introduced to generate diverse unseen class agents, thereby broadening the representation space to cover unknown classes. It is applied at a randomly selected layer in the Vision Transformer, producing smooth low- and high-level representation space boundaries to cover rich attributes. Secondly, it should be noted that unseen class agents generated by the RTM may be mixed with seen class samples. To overcome this challenge, the UCD is designed to generate greater entropy values for unseen classes, thereby distinguishing seen classes from unseen classes. Thirdly, to further mitigate the bias toward seen classes and explore associations between semantics and visual images, the HCR is proposed, which can reconstruct masked pixels based on few discriminative tokens and attribute embeddings. This approach can enable models to have a deep understanding of image contents and build powerful connections between semantic attributes and visual information. Both qualitative and quantitative results demonstrate the effectiveness and usefulness of our proposed M3R model.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot LearningXiangyan Qu, Jing Yu, Keke Gai, Jiamin Zhuang 等ACM MM 2024 · 被引用 5 次
- KNN Transformer with Pyramid Prompts for Few-Shot LearningWenhao Li, Qiangchang Wang, Peng Zhao, Yilong YinACM MM 2024 · 被引用 3 次
- Intra-class Distribution-guided Generative Hashing with Neighbor Refinement for Cross-modal RetrievalHao Sun, Yadong Huo, Qibing Qin, Wenfeng Zhang 等CVPR 2026
- Visual and Semantic Prompt Collaboration for Generalized Zero-Shot LearningHuajie Jiang, Zhengxian Li, Xiaohan Yu, Yongli Hu 等CVPR 2025
相关 Paper
- Progressive Semantic-Guided Vision Transformer for Zero-Shot LearningShiming Chen, Wenjin Hou, Salman H. Khan, Fahad Shahbaz KhanCVPR 2024
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 被引用 43 次
- Meta-Learning for Generalized Zero-Shot LearningVinay Kumar Verma, Dhanajit Brahma, Piyush RaiAAAI 2020 · 被引用 112 次
- TransZero: Attribute-Guided Transformer for Zero-Shot LearningShiming Chen, Ziming Hong, Yang Liu, Guo-Sen Xie 等AAAI 2022 · 被引用 185 次
- SVIP: Semantically Contextualized Visual Patches for Zero-Shot LearningZhi Chen, Zecheng Zhao, Jingcai Guo, Jingjing Li 等ICCV 2025 · 被引用 8 次
