Cross-Modal Relation-Aware Networks for Audio-Visual Event Localization
Haoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan, Chuang Gan
Abstract
We address the challenging task of event localization, which requires the machine to localize an event and recognize its category in unconstrained videos. Most existing methods leverage only the visual information of a video while neglecting its audio information, which, however, can be very helpful and important for event localization. For example, humans often recognize an event by reasoning with the visual and audio content simultaneously. Moreover, the audio information can guide the model to pay more attention on the informative regions of visual scenes, which can help to reduce the interference brought by the background. Motivated by these, in this paper, we propose a relation-aware network to leverage both audio and visual information for accurate event localization. Specifically, to reduce the interference brought by the background, we propose an audio-guided spatial-channel attention module to guide the model to focus on event-relevant visual regions. Besides, we propose to build connections between visual and audio modalities with a relation-aware module. In particular, we learn the representations of video and/or audio segments by aggregating information from the other modality according to the cross-modal relations. Last, relying on the relation-aware representations, we conduct event localization by predicting the event relevant score and classification score. Extensive experimental results demonstrate that our method significantly outperforms the state-of-the-arts in both supervised and weakly-supervised AVE settings. The source code is available at https://github.com/FloretCat/CMRAN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54c4373a-316e-4e7c-863f-b38945b2d957Cited by top-tier papers23
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- Achieving Cross Modal Generalization with Multimodal Unified RepresentationYan Xia, Hai Huang, Jieming Zhu, Zhou ZhaoNeurIPS 2023 · 84 citations
- Cross-modal Background Suppression for Audio-Visual Event LocalizationYan Xia, Zhou ZhaoCVPR 2022 · 62 citations
- MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video ParsingJiashuo Yu, Ying Cheng, Rui-Wei Zhao, Rui Feng et al.ACM MM 2022 · 62 citations
- Cross-modal Prompts: Adapting Large Pre-trained Models for Audio-Visual Downstream TasksHaoyi Duan, Yan Xia, Mingze Zhou, Li Tang et al.NeurIPS 2023 · 59 citations
Builds on10
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 395 citations
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 233 citations
- Location-Aware Graph Convolutional Networks for Video Question AnsweringDeng Huang, Peihao Chen, Runhao Zeng, Qing Du et al.AAAI 2020 · 187 citations
Related papers
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao et al.ACM MM 2024 · 8 citations
- Span-based Audio-Visual LocalizationYiling Wu, Xinfeng Zhang, Yaowei Wang, Qingming HuangACM MM 2022 · 7 citations
- Audio-Visual Semantic Graph Network for Audio-Visual Event LocalizationLiang Liu, Shuaiyong Li, Yongqiang ZhuCVPR 2025
- Learning Event-Specific Localization Preferences for Audio-Visual Event LocalizationShiping Ge, Zhiwei Jiang, Yafeng Yin, Cong Wang et al.ACM MM 2023 · 12 citations
- Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action LocalizationJun-Tae Lee, Mihir Jain, Hyoungwoo Park, Sungrack YunICLR 2021 · 73 citations
