CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-Training
Yuxin Guo, Siyang Sun, Shuailei Ma, Kecheng Zheng, Xiaoyi Bao, Shijie Ma, Wei Zou, Yun Zheng
摘要
Learning joint and coordinated features across modalities is essential for many audio-visual tasks. Existing pretraining methods primarily focus on global information, neglecting fine-grained features and positions, leading to suboptimal performance in dense prediction tasks. To address this issue, we take a further step towards region-aware audio-visual pre-training and propose CrossMAE, which excels in Cross-modality interaction and region alignment. Specifically, we devise two masked autoencoding (MAE) pretext tasks at both pixel and embedding levels, namely Cross-Conditioned Reconstruction and Cross-Embedding Reconstruction. Taking the visual modality as an example (the same goes for audio), in Cross-Conditioned Reconstruction, the visual modality reconstructs the input image pixels conditioned on audio Attentive Tokens. As for the more challenging Cross-Embedding Reconstruction, unmasked visual tokens reconstruct complete audio features under the guidance of Learnable Queries implying positional information, which effectively enhances the interaction between modalities and exploits fine-grained semantics. Experimental results demonstrate that CrossMAE achieves state-of-the-art performance not only in classification and retrieval, but also in dense prediction tasks. Furthermore, we dive into the mechanism of modal interaction and region alignment of CrossMAE, highlighting the effectiveness of the proposed components.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsHaokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui 等NeurIPS 2024 · 被引用 206 次
- Happy: A Debiased Learning Framework for Continual Generalized Category DiscoveryShijie Ma, Fei Zhu, Zhun Zhong, Wenzhuo Liu 等NeurIPS 2024 · 被引用 28 次
- Learning Rate Scaling across LoRA Ranks and Transfer to Full FinetuningNan Chen, Soledad Villar, Soufiane HayouICML 2026 · 被引用 8 次
- PreFM: Online Audio-Visual Event Parsing via Predictive Future ModelingXiao Yu, Yan Fang, Yao Zhao, Yunchao WeiNeurIPS 2025 · 被引用 4 次
- GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric EnhancersShijie Ma, Yuying Ge, Teng Wang, Yuxin Guo 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals 等ICML 2021 · 被引用 1,399 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen 等NeurIPS 2021 · 被引用 884 次
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski 等NeurIPS 2022 · 被引用 524 次
相关 Paper
- Contrastive Audio-Visual Masked AutoencoderYuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath 等ICLR 2023 · 被引用 17 次
- CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained AlignmentEdson Araujo, Andrew Rouditchenko, Yuan Gong, Saurabhchand Bhati 等CVPR 2025
- DropMAE: Masked Autoencoders with Spatial-Attention Dropout for Tracking TasksQiangqiang Wu, Tianyu Yang, Ziquan Liu, Baoyuan Wu 等CVPR 2023
- Stare at What You See: Masked Image Modeling without ReconstructionHongwei Xue, Peng Gao, Hongyang Li, Yu Qiao 等CVPR 2023
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng 等ACM MM 2020 · 被引用 93 次
