CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization
Xiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao, Guobin Shen, Qingqun Kong, Xin Yang, Yi Zeng
摘要
The audio-visual event localization task requires identifying concurrent visual and auditory events from unconstrained videos within a model, locating them, and classifying their category. The efficient extraction and integration of audio and visual modal information have always been challenging in this field. In this paper, we introduce CACE-Net, which differs from most existing methods that solely use audio signals to guide visual information. We propose an audio-visual co-guidance attention mechanism that allows for adaptive bi-directional cross-modal attentional guidance between audio and visual clues, thus reducing inconsistencies between modalities. Moreover, we have observed that existing methods have difficulty distinguishing between similar background and event and lack the fine-grained features for event classification. Consequently, we employ background-event contrast enhancement to increase the discrimination of fused features and fine-tuned pre-trained model to extract more discernible features from complex multimodal inputs. Experiments on the AVE dataset demonstrate that CACE-Net sets a new benchmark in the audio-visual event localization task, proving the effectiveness of our proposed methods in handling complex multimodal learning and event localization in unconstrained videos. Code is available at https://github.com/Brain-Cog-Lab/CACE-Net.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event LocalizationJinxing Zhou, Ziheng Zhou, Yanghao Zhou, Yuxin Mao 等AAAI 2026 · 被引用 4 次
- PreFM: Online Audio-Visual Event Parsing via Predictive Future ModelingXiao Yu, Yan Fang, Yao Zhao, Yunchao WeiNeurIPS 2025 · 被引用 4 次
- Towards Open-Vocabulary Audio-Visual Event LocalizationJinxing Zhou, Dan Guo, Ruohao Guo, Yuxin Mao 等CVPR 2025
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 被引用 233 次
- Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingDi Hu, Rui Qian, Minyue Jiang, Xiao Tan 等NeurIPS 2020 · 被引用 156 次
- Cross-Modal Relation-Aware Networks for Audio-Visual Event LocalizationHaoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan 等ACM MM 2020 · 被引用 97 次
相关 Paper
- Cross-Modal Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationHanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang 等AAAI 2020 · 被引用 110 次
- Cross-modal Background Suppression for Audio-Visual Event LocalizationYan Xia, Zhou ZhaoCVPR 2022 · 被引用 62 次
- Span-based Audio-Visual LocalizationYiling Wu, Xinfeng Zhang, Yaowei Wang, Qingming HuangACM MM 2022 · 被引用 7 次
- Audio-Visual Semantic Graph Network for Audio-Visual Event LocalizationLiang Liu, Shuaiyong Li, Yongqiang ZhuCVPR 2025
- Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action LocalizationJun-Tae Lee, Mihir Jain, Hyoungwoo Park, Sungrack YunICLR 2021 · 被引用 73 次
