Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and Baseline
Tiantian Geng, Teng Wang, Jinming Duan, Runmin Cong, Feng Zheng
摘要
Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual events with different categories. To better adapt to real-life applications, in this paper we focus on the task of dense-localizing audio-visual events, which aims to jointly localize and recognize all audio-visual events occurring in an untrimmed video. The problem is challenging as it requires fine-grained audio-visual scene and context understanding. To tackle this problem, we introduce the first Untrimmed Audio-Visual (UnAV-J 00) dataset, which contains 10K untrimmed videos with over 30K audio-visual events. Each video has 2.8 audio-visual events on average, and the events are usually related to each other and might co-occur as in real-life scenes. Next, we formulate the task using a new learning-based framework, which is capable of fully integrating audio and visual modalities to localize audio-visual events with various lengths and capture dependencies between them in a single pass. Extensive experiments demonstrate the effectiveness of our method as well as the significance of multi-scale cross-modal perception and dependency modeling for this task. The dataset and code are available at https://unav100.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal UnderstandingYunlong Tang, Daiki Shimada, Jing Bi, Mingqian Feng 等AAAI 2025 · 被引用 29 次
- AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMsLidong Lu, Guo Chen, Zhu Wei, Zhiqi Li 等CVPR 2026 · 被引用 23 次
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video HaystacksSanjoy Chowdhury, Mohamed Elmoghany, Yohan Abeysinghe, Junjie Fei 等NeurIPS 2025 · 被引用 14 次
- Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLMZinuo Li, Xian Zhang, Yongxin Guo, Mohammed Bennamoun 等NeurIPS 2025 · 被引用 9 次
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao 等ACM MM 2024 · 被引用 8 次
它引用的顶会 Paper15
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen 等NeurIPS 2021 · 被引用 884 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 被引用 233 次
- Relaxed Transformer Decoders for Direct Action Proposal GenerationJing Tan, Jiaqi Tang, Limin Wang, Gangshan WuICCV 2021 · 被引用 220 次
- Video Self-Stitching Graph Network for Temporal Action LocalizationChen Zhao, Ali K. Thabet, Bernard GhanemICCV 2021 · 被引用 179 次
相关 Paper
- Dense Audio-Visual Event Localization Under Cross-Modal Consistency and Multi-Temporal Granularity CollaborationZiheng Zhou, Jinxing Zhou, Wei Qian, Shengeng Tang 等AAAI 2025 · 被引用 4 次
- CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event LocalizationJinxing Zhou, Ziheng Zhou, Yanghao Zhou, Yuxin Mao 等AAAI 2026 · 被引用 4 次
- Benchmarking Audio Visual Segmentation for Long-Untrimmed VideosChen Liu, Peike Patrick Li, Qingtao Yu, Hongwei Sheng 等CVPR 2024
- Towards Open-Vocabulary Audio-Visual Event LocalizationJinxing Zhou, Dan Guo, Ruohao Guo, Yuxin Mao 等CVPR 2025
- Learning Event-Specific Localization Preferences for Audio-Visual Event LocalizationShiping Ge, Zhiwei Jiang, Yafeng Yin, Cong Wang 等ACM MM 2023 · 被引用 12 次
