Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
Pengcheng Zhao, Jinxing Zhou, Yang Zhao, Dan Guo, Yanxiang Chen
Abstract
The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works directly utilize the extracted holistic audio and visual features for intra- and cross-modal temporal interactions. However, each segment may contain multiple events, resulting in semantically mixed holistic features that can lead to semantic interference during intra- or cross-modal interactions: the event semantics of one segment may incorporate semantics of unrelated events from other segments. To address this issue, our method begins with a Class-Aware Feature Decoupling (CAFD) module, which explicitly decouples the semantically mixed features into distinct class-wise features, including multiple event-specific features and a dedicated background feature. The decoupled class-wise features enable our model to selectively aggregate useful semantics for each segment from clearly matched classes contained in other segments, preventing semantic interference from irrelevant classes. Specifically, we further design a Fine-Grained Semantic Enhancement module for encoding intra- and cross-modal relations. It comprises a Segment-wise Event Co-occurrence Modeling (SECM) block and a Local-Global Semantic Fusion (LGSF) block. The SECM exploits inter-class dependencies of concurrent events within the same timestamp with the aid of a novel event co-occurrence loss. The LGSF further enhances the event semantics of each segment by incorporating relevant semantics from more informative global video features. Extensive experiments validate the effectiveness of the proposed modules and loss functions, resulting in a new state-of-the-art parsing performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual SegmentationJinxing Zhou, Yanghao Zhou, Mingfei Han, Tong Wang et al.AAAI 2026 · 5 citations
- CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event LocalizationJinxing Zhou, Ziheng Zhou, Yanghao Zhou, Yuxin Mao et al.AAAI 2026 · 4 citations
- Dense Audio-Visual Event Localization Under Cross-Modal Consistency and Multi-Temporal Granularity CollaborationZiheng Zhou, Jinxing Zhou, Wei Qian, Shengeng Tang et al.AAAI 2025 · 4 citations
- PreFM: Online Audio-Visual Event Parsing via Predictive Future ModelingXiao Yu, Yan Fang, Yao Zhao, Yunchao WeiNeurIPS 2025 · 4 citations
- Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment LocalizationCailing Han, Zhangbin Li, Jinxing Zhou, Wei Qian et al.CVPR 2026 · 1 citation
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingYan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin et al.NeurIPS 2021 · 94 citations
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng et al.ACM MM 2020 · 93 citations
- Multi-modal Grouping Network for Weakly-Supervised Audio-Visual Video ParsingShentong Mo, Yapeng TianNeurIPS 2022 · 73 citations
- MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video ParsingJiashuo Yu, Ying Cheng, Rui-Wei Zhao, Rui Feng et al.ACM MM 2022 · 62 citations
Related papers
- Multi-Modal and Multi-Scale Temporal Fusion Architecture Search for Audio-Visual Video ParsingJiayi Zhang, Weixin LiACM MM 2023 · 5 citations
- Exploring Heterogeneous Clues for Weakly-Supervised Audio-Visual Video ParsingYu Wu, Yi YangCVPR 2021
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao et al.ACM MM 2024 · 8 citations
- UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video ParsingYung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang, François G. Germain et al.CVPR 2025
- Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal InteractionMingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li et al.AAAI 2026 · 1 citation
