Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action Localization
Jun-Tae Lee, Mihir Jain, Hyoungwoo Park, Sungrack Yun
Abstract
Temporally localizing actions in videos is one of the key components for video understanding. Learning from weakly-labeled data is seen as a potential solution towards avoiding expensive frame-level annotations. Different from other works which only depend on visual-modality, we propose to learn richer audiovisual representation for weakly-supervised action localization. First, we propose a multi-stage cross-attention mechanism to collaboratively fuse audio and visual features, which preserves the intra-modal characteristics. Second, to model both foreground and background frames, we construct an open-max classifier that treats the background class as an open-set. Third, for precise action localization, we design consistency losses to enforce temporal continuity for the action class prediction, and also help with foreground-prediction reliability. Extensive experiments on two publicly available video-datasets (AVE and ActivityNet1.2) show that the proposed method effectively fuses audio and visual modalities, and achieves the state-of-the-art results for weakly-supervised action localization.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 91c7cccc-1b29-46fd-9376-bd7018377d3bCited by top-tier papers16
- Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingYan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin et al.NeurIPS 2021 · 94 citations
- Learning Action Completeness from Points for Weakly-supervised Temporal Action LocalizationPilhyeon Lee, Hyeran ByunICCV 2021 · 81 citations
- CATN: Cross Attentive Tree-Aware Network for Multivariate Time Series ForecastingHui He, Qi Zhang, Simeng Bai, Kun Yi et al.AAAI 2022 · 55 citations
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 37 citations
- Audio-Adaptive Activity Recognition Across Video DomainsYunhua Zhang, Hazel Doughty, Ling Shao, Cees G. M. SnoekCVPR 2022 · 31 citations
Related papers
- Foreground-Action Consistency Network for Weakly Supervised Temporal Action LocalizationLinjiang Huang, Liang Wang, Hongsheng LiICCV 2021 · 91 citations
- Span-based Audio-Visual LocalizationYiling Wu, Xinfeng Zhang, Yaowei Wang, Qingming HuangACM MM 2022 · 7 citations
- Weakly-Supervised Action Localization by Generative Attention ModelingBaifeng Shi, Qi Dai, Yadong Mu, Jingdong WangCVPR 2020
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao et al.ACM MM 2024 · 8 citations
- Cross-Modal Relation-Aware Networks for Audio-Visual Event LocalizationHaoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan et al.ACM MM 2020 · 97 citations
