Multi-Modal Contrastive Masked Autoencoders: A Two-Stage Progressive Pre-training Approach for RGBD Datasets
Muhammad Abdullah Jamal, Omid Mohareri
摘要
In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Autoencoder and Denoising techniques. Our proposed approach consists of two stages. In the first stage, we pre-train the model using contrastive learning to learn cross-modal representations. In the second stage, we further pre-train the model using masked autoencoding and denoising/noise prediction used in diffusion models. Masked autoencoding focuses on reconstructing the missing patches in the input modality using local spatial correlations, while denoising learns high frequency components of the input data. Moreover, it incorporates global distillation in the second stage by leveraging the knowledge acquired in stage one. Our approach is scalable, robust and suitable for pre-training RGB-D datasets. Extensive experiments on multiple datasets such as ScanNet, NYUv2 and SUN RGB-D show the efficacy and superior performance of our approach. Specifically, we show an improvement of +1.3% mIoU against Mask3D on ScanNet semantic segmentation. We further demonstrate the effectiveness of our approach in low-data regime by evaluating it for semantic segmentation task against the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Calibrated Multimodal Representation Learning with Missing ModalitiesXiaohao Liu, Xiaobo Xia, Jiaheng Wei, Shuo Yang 等ICML 2026 · 被引用 5 次
- A Mixed Diet Makes DINO An Omnivorous Vision EncoderRishabh Kabra, Maks Ovsjanikov, Drew A. Hudson, Ye Xia 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- CoMAE: Single Model Hybrid Pre-training on Small-Scale RGB-D DatasetsJiange Yang, Sheng Guo, Gangshan Wu, Limin WangAAAI 2023 · 被引用 12 次
- PiMAE: Point Cloud and Image Interactive Masked Autoencoders for 3D Object DetectionAnthony Chen, Kevin Zhang, Renrui Zhang, Zihan Wang 等CVPR 2023
- MaskMentor: Unlocking the Potential of Masked Self-Teaching for Missing Modality RGB-D Semantic SegmentationZhida Zhao, Jia Li, Lijun Wang, Yifan Wang 等ACM MM 2024 · 被引用 1 次
- Multi-modal Vision Pre-training for Medical Image AnalysisShaohao Rui, Lingzhi Chen, Zhenyu Tang, Lilong Wang 等CVPR 2025
- Progressive Mask Distillation for Self-supervised Video RepresentationKewei Wu, Chong Liang, Zhao Xie, Dan GuoCVPR 2026
