M3R: Masked Token Mixup and Cross-Modal Reconstruction for Zero-Shot Learning
Peng Zhao, Qiangchang Wang, Yilong Yin
Abstract
In the zero-shot learning (ZSL), learned representation spaces are often biased toward seen classes, thus limiting the ability to predict previously unseen classes. In this paper, we propose Masked token Mixup and cross-Modal Reconstruction for zero-shot learning, termed as M3R, which can significantly alleviate the bias toward seen classes. The M3R mainly consists of Random Token Mixup (RTM), Unseen Class Detection (UCD), and Hard Cross-modal Reconstruction (HCR). Firstly, mappings without proper adaptations to unseen classes would cause the bias toward seen classes. To address this issue, the RTM is introduced to generate diverse unseen class agents, thereby broadening the representation space to cover unknown classes. It is applied at a randomly selected layer in the Vision Transformer, producing smooth low- and high-level representation space boundaries to cover rich attributes. Secondly, it should be noted that unseen class agents generated by the RTM may be mixed with seen class samples. To overcome this challenge, the UCD is designed to generate greater entropy values for unseen classes, thereby distinguishing seen classes from unseen classes. Thirdly, to further mitigate the bias toward seen classes and explore associations between semantics and visual images, the HCR is proposed, which can reconstruct masked pixels based on few discriminative tokens and attribute embeddings. This approach can enable models to have a deep understanding of image contents and build powerful connections between semantic attributes and visual information. Both qualitative and quantitative results demonstrate the effectiveness and usefulness of our proposed M3R model.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 38ca2c6e-8489-417a-bce2-c4d74be858f6Cited by top-tier papers4
- Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot LearningXiangyan Qu, Jing Yu, Keke Gai, Jiamin Zhuang et al.ACM MM 2024 · 5 citations
- KNN Transformer with Pyramid Prompts for Few-Shot LearningWenhao Li, Qiangchang Wang, Peng Zhao, Yilong YinACM MM 2024 · 3 citations
- Intra-class Distribution-guided Generative Hashing with Neighbor Refinement for Cross-modal RetrievalHao Sun, Yadong Huo, Qibing Qin, Wenfeng Zhang et al.CVPR 2026
- Visual and Semantic Prompt Collaboration for Generalized Zero-Shot LearningHuajie Jiang, Zhengxian Li, Xiaohan Yu, Yongli Hu et al.CVPR 2025
Related papers
- Progressive Semantic-Guided Vision Transformer for Zero-Shot LearningShiming Chen, Wenjin Hou, Salman H. Khan, Fahad Shahbaz KhanCVPR 2024
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 43 citations
- Meta-Learning for Generalized Zero-Shot LearningVinay Kumar Verma, Dhanajit Brahma, Piyush RaiAAAI 2020 · 112 citations
- TransZero: Attribute-Guided Transformer for Zero-Shot LearningShiming Chen, Ziming Hong, Yang Liu, Guo-Sen Xie et al.AAAI 2022 · 185 citations
- SVIP: Semantically Contextualized Visual Patches for Zero-Shot LearningZhi Chen, Zecheng Zhao, Jingcai Guo, Jingjing Li et al.ICCV 2025 · 8 citations
