Multi-modal Grouping Network for Weakly-Supervised Audio-Visual Video Parsing
Shentong Mo, Yapeng Tian
摘要
The audio-visual video parsing task aims to parse a video into modality-and category-aware temporal segments. Previous work mainly focuses on weakly-supervised approaches, which learn from video-level event labels. During training, they do not know which modality perceives and meanwhile which temporal segment contains the video event. Since there is no explicit grouping in the existing frameworks, the modality and temporal uncertainties make these meth-ods suffer from false predictions. For instance, segments in the same category could be predicted in different event classes. Learning compact and discriminative multi-modal subspaces is essential for mitigating the issue. To this end, in this paper, we propose a novel Multi-modal Grouping Network, namely MGN, for explicitly semantic-aware grouping. Specifically, MGN aggregates event-aware unimodal features through unimodal grouping in terms of learnable categorical embedding tokens. Furthermore, it leverages the cross-modal grouping for modality-aware prediction to match the video-level target. Our simple framework achieves improving results against previous baselines on weakly-supervised audio-visual video parsing. In addition, our MGN is much more lightweight, using only 47.2% of the parameters of baselines (17 MB vs. 36 MB). Code is available at https://github.com/stoneMo/MGN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Cross-modal Prompts: Adapting Large Pre-trained Models for Audio-Visual Downstream TasksHaoyi Duan, Yan Xia, Mingze Zhou, Li Tang 等NeurIPS 2023 · 被引用 59 次
- Audio-Visual Class-Incremental LearningWeiguo Pian, Shentong Mo, Yunhui Guo, Yapeng TianICCV 2023 · 被引用 44 次
- Class-Incremental Grouping Network for Continual Audio-Visual LearningShentong Mo, Weiguo Pian, Yapeng TianICCV 2023 · 被引用 34 次
- Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event ParserYung-Hsuan Lai, Yen-Chun Chen, Frank WangNeurIPS 2023 · 被引用 27 次
- A Unified Audio-Visual Learning Framework for Localization, Separation, and RecognitionShentong Mo, Pedro MorgadoICML 2023 · 被引用 27 次
它引用的顶会 Paper13
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 被引用 233 次
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- Learning Representations from Audio-Visual Spatial AlignmentPedro Morgado, Yi Li, Nuno VasconcelosNeurIPS 2020 · 被引用 149 次
- Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingYan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin 等NeurIPS 2021 · 被引用 94 次
相关 Paper
- DHHN: Dual Hierarchical Hybrid Network for Weakly-Supervised Audio-Visual Video ParsingXun Jiang, Xing Xu, Zhiguo Chen, Jingran Zhang 等ACM MM 2022 · 被引用 35 次
- UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video ParsingYung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang, François G. Germain 等CVPR 2025
- Exploring Heterogeneous Clues for Weakly-Supervised Audio-Visual Video ParsingYu Wu, Yi YangCVPR 2021
- MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video ParsingLangyu Wang, Bingke Zhu, Yingying Chen, Yiyuan Zhang 等ICCV 2025 · 被引用 1 次
- Audio-Visual Grouping Network for Sound Localization from MixturesShentong Mo, Yapeng TianCVPR 2023
