CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
Edson Araujo, Andrew Rouditchenko, Yuan Gong, Saurabhchand Bhati, Samuel Thomas, Brian Kingsbury, Leonid Karlinsky, Rogério Feris, James R. Glass, Hilde Kuehne
摘要
Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained temporal correspondences with visual frames. Additionally, existing methods often struggle with conflicting optimization objectives when trying to jointly learn reconstruction and cross-modal alignment. In this work, we propose CAV-MAE Sync as a simple yet effective extension of the original CAV-MAE [14] framework for self-supervised audiovisual learning. We address three key challenges: First, we tackle the granularity mismatch between modalities by treating audio as a temporal sequence aligned with video frames, rather than using global representations. Second, we resolve conflicting optimization goals by separating contrastive and reconstruction objectives through dedicated global tokens. Third, we improve spatial localization by introducing learnable register tokens that reduce the semantic load on patch tokens. We evaluate the proposed approach on AudioSet, VGG Sound, and the ADE20K Sound dataset on zero-shot retrieval, classification, and localization tasks demonstrating state-of-the-art performance and outperforming more complex architectures. Code is available at https://github.com/edsonroteia/cav-mae-sync.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Hear What Matters! Text-conditioned Selective Video-to-Audio GenerationJunwon Lee, Juhan Nam, Jiyoung LeeCVPR 2026 · 被引用 4 次
- Can Diffusion Models Disentangle? A Theoretical PerspectiveLiming Wang, Muhammad Jehanzeb Mirza, Yishu Gong, Yuan Gong 等NeurIPS 2025 · 被引用 4 次
- Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake LocalizationJiayu Xiong, Jing Wang, Qi Zhang, Wanlong Wang 等CVPR 2026
- AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint ControlXinyue Guo, Xiaoran Yang, Lipan Zhang, Jianxuan Yang 等AAAI 2026
- Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation LearningLinge Wang, Yingying Chen, Bingke Zhu, Lu Zhou 等CVPR 2026
它引用的顶会 Paper25
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan 等ICLR 2024 · 被引用 403 次
- Broaden Your Views for Self-Supervised Video LearningAdrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang 等ICCV 2021 · 被引用 139 次
- Active Contrastive Learning of Audio-Visual Video RepresentationsShuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale SongICLR 2021 · 被引用 109 次
相关 Paper
- Contrastive Audio-Visual Masked AutoencoderYuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath 等ICLR 2023 · 被引用 17 次
- Bio-Inspired Audiovisual Multi-Representation Integration via Self-Supervised LearningZhaojian Li, Bin Zhao, Yuan YuanACM MM 2023 · 被引用 3 次
- Learning Representations from Audio-Visual Spatial AlignmentPedro Morgado, Yi Li, Nuno VasconcelosNeurIPS 2020 · 被引用 149 次
- CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-TrainingYuxin Guo, Siyang Sun, Shuailei Ma, Kecheng Zheng 等CVPR 2024
- Contrastive Learning of Global and Local Video RepresentationsShuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale SongNeurIPS 2021 · 被引用 45 次
