Weakly-Supervised Spoken Video Grounding via Semantic Interaction Learning
Ye Wang, Wang Lin, Shengyu Zhang, Tao Jin, Linjun Li, Xize Cheng, Zhou Zhao
Abstract
The task of spoken video grounding aims to localize moments in videos that are relevant to descriptive spoken queries. However, extracting semantic information from speech and modeling the cross-modal correlation pose two critical challenges. Previous studies solve them by representing spoken queries based on the matched video frames, which require tremendous effort for frame-level labeling. In this work, we investigate weakly-supervised spoken video grounding, i.e., learning to localize moments without expensive temporal annotations. To effectively represent the cross-modal semantics, we propose Semantic Interaction Learning (SIL), a novel framework consisting of the acoustic-semantic pre-training (ASP) and acoustic-visual contrastive learning (AVCL). In ASP, we pre-train an effective encoder for the grounding task with three comprehensive tasks, where the robustness task enhances stability by explicitly capturing the invariance between time-and frequency-domain features, the conciseness task avoids over-smooth attention by compressing long sequence into segments, and the semantic task improves spoken language understanding by modeling the precise semantics. In AVCL, we mine pseudo labels with discriminative sampling strategies and directly strengthen the interaction between speech and video by maximizing their mutual information. Extensive experiments demonstrate the effectiveness and superiority of our method. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d245ed12-c35a-4601-9cb2-7d0b6dd8676fCited by top-tier papers6
- Towards Unified Multimodal Editing with Enhanced Knowledge CollaborationKaihang Pan, Zhaoyu Fan, Juncheng Li, Qifan Yu et al.NeurIPS 2024 · 27 citations
- Exploring Group Video Captioning with Efficient Relational ApproximationWang Lin, Tao Jin, Ye Wang, Wenwen Pan et al.ICCV 2023 · 17 citations
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based AgentsTao Wu, Jingyuan Chen, Wang Lin, Mengze Li et al.ACL 2025 · 16 citations
- Low-rank Prompt Interaction for Continual Vision-Language RetrievalWeicai Yan, Ye Wang, Wang Lin, Zirun Guo et al.ACM MM 2024 · 8 citations
- WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed BenchmarkWang Lin, Feng Wang, Majun Zhang, Wentao Hu et al.ICLR 2026 · 2 citations
Builds on19
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Self-Supervised Contrastive Pre-Training For Time Series via Time-Frequency ConsistencyXiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, Marinka ZitnikNeurIPS 2022 · 558 citations
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 279 citations
Related papers
- Counterfactual Contrastive Learning for Weakly-Supervised Vision-Language GroundingZhu Zhang, Zhou Zhao, Zhijie Lin, Jieming Zhu et al.NeurIPS 2020 · 74 citations
- Video-Guided Curriculum Learning for Spoken Video GroundingYan Xia, Zhou Zhao, Shangwei Ye, Yang Zhao et al.ACM MM 2022 · 7 citations
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationPeijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er et al.AAAI 2023 · 13 citations
- Learning Multi-Scale Video-Text Correspondence for Weakly Supervised Temporal Article GrondingWenjia Geng, Yong Liu, Lei Chen, Sujia Wang et al.AAAI 2024 · 3 citations
- D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance AnnotationHanjun Li, Xiujun Shu, Sunan He, Ruizhi Qiao et al.ICCV 2023 · 21 citations
