Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence Localization
Minghang Zheng, Shaogang Gong, Hailin Jin, Yuxin Peng, Yang Liu
Abstract
Video sentence localization aims to locate moments in an unstructured video according to a given natural language query. A main challenge is the expensive annotation costs and the annotation bias. In this work, we study video sentence localization in a zero-shot setting, which learns with only video data without any annotation. Existing zero-shot pipelines usually generate event proposals and then generate a pseudo query for each event proposal. However, their event proposals are obtained via visual feature clustering, which is query-independent and inaccurate; and the pseudo-queries are short or less interpretable. Moreover, existing approaches ignores the risk of pseudo-label noise when leveraging them in training. To address the above problems, we propose a Structure-based Pseudo Label generation (SPL), which first generate free-form interpretable pseudo queries before constructing query-dependent event proposals by modeling the event temporal structure. To mitigate the effect of pseudo-label noise, we propose a noise-resistant iterative method that repeatedly re-weight the training sample based on noise estimation to train a grounding model and correct pseudo labels. Experiments on the ActivityNet Captions and Charades-STA datasets demonstrate the advantages of our approach. Code can be found at https://github.com/minghangz/SPL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f763451-a2b7-4b68-9c0c-1f5958e740aaCited by top-tier papers9
- Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in VideoZhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang et al.AAAI 2024 · 17 citations
- PlanLLM: Video Procedure Planning with Refinable Large Language ModelsDejie Yang, Zijing Zhao, Yang LiuAAAI 2025 · 8 citations
- Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept CalibrationTing Lei, Shaofeng Yin, Qingchao Chen, Yuxin Peng et al.ICCV 2025 · 6 citations
- ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual GroundingMinghang Zheng, Jiahua Zhang, Qingchao Chen, Yuxin Peng et al.ACM MM 2024 · 5 citations
- Generative Video Diffusion for Unseen Novel Semantic Video Moment RetrievalDezhao Luo, Shaogang Gong, Jiabo Huang, Hailin Jin et al.AAAI 2025 · 4 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang et al.AAAI 2020 · 170 citations
- Memory-Guided Semantic Learning Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Xing Di, Yu Cheng et al.AAAI 2022 · 83 citations
- Cross-Sentence Temporal and Semantic Relations in Video Activity LocalisationJiabo Huang, Yang Liu, Shaogang Gong, Hailin JinICCV 2021 · 77 citations
Related papers
- Zero-shot Natural Language Video LocalizationJinwoo Nam, Daechul Ahn, Dongyeop Kang, Seong Jong Ha et al.ICCV 2021 · 60 citations
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng et al.CVPR 2022 · 108 citations
- Unsupervised Temporal Video Grounding with Deep Semantic ClusteringDaizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di et al.AAAI 2022 · 52 citations
- Prompt-based Zero-shot Video Moment RetrievalGuolong Wang, Xun Wu, Zhaoyuan Liu, Junchi YanACM MM 2022 · 33 citations
- Boundary-Aware Temporal Dynamic Pseudo-Supervision Pairs Generation for Zero-Shot Natural Language Video LocalizationXiongwen Deng, Haoyu Tang, Han Jiang, Qinghai Zheng et al.AAAI 2025
