Towards Open-Vocabulary Audio-Visual Event Localization
Jinxing Zhou, Dan Guo, Ruohao Guo, Yuxin Mao, Jingjing Hu, Yiran Zhong, Xiaojun Chang, Meng Wang
Abstract
The Audio-Visual Event Localization (AVEL) task aims to temporally locate and classify video events that are both audible and visible. Most research in this field assumes a closed-set setting, which restricts these models' ability to handle test data containing event categories absent (unseen) during training. Recently, a few studies have explored AVEL in an open-set setting, enabling the recognition of unseen events as "unknown", but without providing category-specific semantics. In this paper, we advance the field by introducing the Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) problem, which requires localizing audio-visual events and predicting explicit categories for both seen and unseen test data at inference. To address this new task, we propose the OV-AVEBench dataset, comprising 24,800 videos across 67 real-life audiovisual scenes (seen:unseen = 46:21), each with manual segment-level annotation. We also establish three evaluation metrics for this task. Moreover, we investigate two baseline approaches, one training-free and one using a further fine-tuning paradigm. Specifically, we utilize the unified multimodal space from the pretrained ImageBind model to extract audio, visual, and textual (event classes) features. The training-free baseline then determines predictions by comparing the consistency of audio-text and visualtext feature similarities. The fine-tuning baseline incorporates lightweight temporal layers to encode temporal relations within the audio and visual modalities, using OV-AVEBench training data for model fine-tuning. We evaluate these baselines on the proposed OV-AVEBench dataset and discuss potential directions for future work in this new field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78dc2694-6952-480d-b12d-01ac8e422172Cited by top-tier papers12
- Patch-level Sounding Object Tracking for Audio-Visual Question AnsweringZhangbin Li, Jinxing Zhou, Jing Zhang, Shengeng Tang et al.AAAI 2025 · 20 citations
- Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video ParsingPengcheng Zhao, Jinxing Zhou, Yang Zhao, Dan Guo et al.AAAI 2025 · 19 citations
- Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual SegmentationJinxing Zhou, Yanghao Zhou, Mingfei Han, Tong Wang et al.AAAI 2026 · 5 citations
- EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound GenerationBingxuan Li, Yiming Cui, Yicheng He, Yiwei Wang et al.CVPR 2026 · 5 citations
- CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event LocalizationJinxing Zhou, Ziheng Zhou, Yanghao Zhou, Yuxin Mao et al.AAAI 2026 · 4 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 233 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
- Cross-Modal Relation-Aware Networks for Audio-Visual Event LocalizationHaoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan et al.ACM MM 2020 · 97 citations
- Cross-modal Background Suppression for Audio-Visual Event LocalizationYan Xia, Zhou ZhaoCVPR 2022 · 62 citations
Related papers
- OV-DAVEL: Towards Open-Vocabulary Dense Audio-Visual Event Localization in Untrimmed VideosJiale Yu, Baopeng Zhang, Zhu Teng, Jianping FanACM MM 2025
- Open-Vocabulary Audio-Visual Semantic SegmentationRuohao Guo, Liao Qu, Dantong Niu, Yanyu Qi et al.ACM MM 2024 · 4 citations
- OpenAVE: Moving towards Open Set Audio-Visual Event LocalizationJiale Yu, Baopeng Zhang, Zhu Teng, Jianping FanACM MM 2024 · 2 citations
- Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and BaselineTiantian Geng, Teng Wang, Jinming Duan, Runmin Cong et al.CVPR 2023
- Learning Event-Specific Localization Preferences for Audio-Visual Event LocalizationShiping Ge, Zhiwei Jiang, Yafeng Yin, Cong Wang et al.ACM MM 2023 · 12 citations
