Open-Set Cross Modal Generalization via Multimodal Unified Representation
Hai Huang, Yan Xia, Shulei Wang, Hanting Wang, Minghui Fang, Shengpeng Ji, Sashuai Zhou, Tao Jin, Zhou Zhao
摘要
This paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates multimodal unified representations in open-set conditions, addressing the limitations of prior closed-set cross-modal evaluations. OSCMG requires not only cross-modal knowledge transfer but also robust generalization to unseen classes within new modalities, a scenario frequently encountered in real-world applications. Existing multimodal unified representation work lacks consideration for open-set environments. To tackle this, we propose MICU, comprising two key components: Fine-Coarse Masked multimodal InfoNCE (FCMI) and Cross modal Unified Jigsaw Puzzles (CUJP). FCMI enhances multimodal alignment by applying contrastive learning at both holistic semantic and temporal levels, incorporating masking to enhance generalization. CUJP enhances feature diversity and model uncertainty by integrating modality-agnostic feature selection with self-supervised learning, thereby strengthening the model's ability to handle unknown categories in open-set tasks. Extensive experiments on CMG and the newly proposed OSCMG validate the effectiveness of our approach. The code is available at https://github.com/haihuangcode/CMG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Chat-Driven Text Generation and Interaction for Person RetrievalZequn Xie, Chuxin Wang, Yeqiang Wang, Sihang Cai 等EMNLP 2025 · 被引用 12 次
- TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather RemovalHanting Wang, Shengpeng Ji, Shulei Wang, Hai Huang 等ACM MM 2025
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Domain Generalization for Medical Imaging Classification with Linear-Dependency RegularizationHaoliang Li, Yufei Wang, Renjie Wan, Shiqi Wang 等NeurIPS 2020 · 被引用 233 次
- UNIFIED-IO: A Unified Model for Vision, Language, and Multi-modal TasksJiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi 等ICLR 2023 · 被引用 110 次
相关 Paper
- Achieving Cross Modal Generalization with Multimodal Unified RepresentationYan Xia, Hai Huang, Jieming Zhu, Zhou ZhaoNeurIPS 2023 · 被引用 84 次
- Contrastive Multimodal Fusion with TupleInfoNCEYunze Liu, Qingnan Fan, Shanghang Zhang, Hao Dong 等ICCV 2021 · 被引用 84 次
- The Finer the Better: Towards Granular-aware Open-set Domain GeneralizationYunyun Wang, Zheng Duan, Xinyue Liao, Ke-Jia Chen 等AAAI 2026
- Towards Out-of-Modal Generalization without Instance-level Modal CorrespondenceZhuo Huang, Gang Niu, Bo Han, Masashi Sugiyama 等ICLR 2025
- InfMasking: Unleashing Synergistic Information by Contrastive Multimodal InteractionsLiangjian Wen, Qun Dai, Jianzhuang Liu, Jiangtao Zheng 等NeurIPS 2025 · 被引用 10 次
