MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model
Yatai Ji, Junjie Wang, Yuan Gong, Lin Zhang, Yanru Zhu, Hongfa Wang, Jiaxing Zhang, Tetsuya Sakai, Yujiu Yang
Abstract
Multimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets. Such uncertainty is problematic for our interpretation, including inter-and intra-modal uncertainty. Little effort has studied the modeling of this uncertainty, particularly in pre-training on unlabeled datasets and fine-tuning in task-specific downstream datasets. In this paper, we project the representations of all modalities as probabilistic distributions via a Probability Distribution Encoder (PDE) by utilizing sequence-level interactions. Compared to the existing deterministic methods, such uncertainty modeling can convey richer multimodal semantic information and more complex relationships. Furthermore, we integrate uncertainty modeling with popular pretraining frameworks and propose suitable pre-training tasks: Distribution-based Vision-Language Contrastive learning (D-VLC), Distribution-based Masked Language Modeling (D-MLM), and Distribution-based Image-Text Matching (D-ITM) . The fine-tuned models are applied to challenging downstream tasks, including image-text retrieval, visual question answering, visual reasoning, and visual entailment, and achieve state-of-the-art results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Improved Probabilistic Image-Text RepresentationsSanghyuk ChunICLR 2024 · 48 citations
- Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identificationZhiwei Zhao, Bin Liu, Yan Lu, Qi Chu et al.AAAI 2024 · 40 citations
- Tackling Uncertain Correspondences for Multi-Modal Entity AlignmentLiyi Chen, Ying Sun, Shengzhe Zhang, Yuyang Ye et al.NeurIPS 2024 · 20 citations
- Diffusion-Inspired Truncated Sampler for Text-Video RetrievalJiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan et al.NeurIPS 2024 · 16 citations
- ExACT: Language-Guided Conceptual Reasoning and Uncertainty Estimation for Event-Based Action Recognition and MoreJiazhou Zhou, Xu Zheng, Yuanhuiyi Lyu, Lin WangCVPR 2024 · 15 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
Related papers
- Seeing What You Miss: Vision-Language Pre-training with Semantic Completion LearningYatai Ji, Rongcheng Tu, Jie Jiang, Weijie Kong et al.CVPR 2023
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningWei Li, Can Gao, Guocheng Niu, Xinyan Xiao et al.ACL 2021
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-trainingHongwei Xue, Yupan Huang, Bei Liu, Houwen Peng et al.NeurIPS 2021 · 100 citations
- MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal EmbeddingsHaonan Chen, Hong Liu, Yuping Luo, Liang Wang et al.ACL 2026 · 20 citations
