Video-Guided Curriculum Learning for Spoken Video Grounding
Yan Xia, Zhou Zhao, Shangwei Ye, Yang Zhao, Haoyuan Li, Yi Ren
摘要
In this paper, we introduce a new task, spoken video grounding (SVG), which aims to localize the desired video fragments from spoken language descriptions. Compared with using text, employing audio requires the model to directly exploit the useful phonemes and syllables related to the video from raw speech. Moreover, we randomly add environmental noises to this speech audio, further increasing the difficulty of this task and better simulating real applications. To rectify the discriminative phonemes and extract video-related information from noisy audio, we develop a novel video-guided curriculum learning (VGCL) during the audio pretraining process, which can make use of the vital visual perceptions to help understand the spoken language and suppress the external noise. Considering during inference the model can not obtain ground truth video segments, we design a curriculum strategy that gradually shifts the input video from the ground truth to the entire video content during pre-training. Finally, the model can learn how to extract critical visual information from the entire video clip to help understand the spoken language. In addition, we collect the first large-scale spoken video grounding dataset based on ActivityNet, which is named as ActivityNet Speech dataset. Extensive experiments demonstrate our proposed video-guided curriculum learning can facilitate the pre-training process to obtain a mutual audio encoder, significantly promoting the performance of spoken video grounding tasks. Moreover, we prove that in the case of noisy sound, our model outperforms the method that grounding video with ASR transcripts, further demonstrating the effectiveness of our curriculum strategy. The code is available at https://github.com/marmot-xy/Spoken-Video-Grounding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- TranSpeech: Speech-to-Speech Translation With Bilateral PerturbationRongjie Huang, Jinglin Liu, Huadai Liu, Yi Ren 等ICLR 2023 · 被引用 17 次
- Weakly-Supervised Spoken Video Grounding via Semantic Interaction LearningYe Wang, Wang Lin, Shengyu Zhang, Tao Jin 等ACL 2023 · 被引用 7 次
- Gloss Attention for Gloss-free Sign Language TranslationAoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin 等CVPR 2023
- DATE: Domain Adaptive Product Seeker for E-CommerceHaoyuan Li, Hao Jiang, Tao Jin, Mengyan Li 等CVPR 2023
它引用的顶会 Paper8
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 被引用 279 次
- Curriculum By SmoothingSamarth Sinha, Animesh Garg, Hugo LarochelleNeurIPS 2020 · 被引用 95 次
- Fine-grained Iterative Attention Network for Temporal Language Localization in VideosXiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng 等ACM MM 2020 · 被引用 92 次
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded SpeechDavid Harwath, Wei-Ning Hsu, James R. GlassICLR 2020 · 被引用 88 次
- Local-Global Video-Text Interactions for Temporal GroundingJonghwan Mun, Minsu Cho, Bohyung HanCVPR 2020
相关 Paper
- Video Object Grounding Using Semantic Roles in Language DescriptionArka Sadhu, Kan Chen, Ram NevatiaCVPR 2020
- Curriculum Multi-Negative Augmentation for Debiased Video GroundingXiaohan Lan, Yitian Yuan, Hong Chen, Xin Wang 等AAAI 2023 · 被引用 26 次
- Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form SentencesZhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang 等CVPR 2020
- SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video GroundingHong Gao, Xiangkai Xu, Bin Zhong, Junjie Yin 等CVPR 2026
- Learning Transferable Spatiotemporal Representations from Natural Script KnowledgeZiyun Zeng, Yuying Ge, Xihui Liu, Bin Chen 等CVPR 2023
