WINNER: Weakly-supervised hIerarchical decompositioN and aligNment for spatio-tEmporal video gRounding
Mengze Li, Han Wang, Wenqiao Zhang, Jiaxu Miao, Zhou Zhao, Shengyu Zhang, Wei Ji, Fei Wu
Abstract
Spatio-temporal video grounding aims to localize the aligned visual tube corresponding to a language query. Existing techniques achieve such alignment by exploiting dense boundary and bounding box annotations, which can be prohibitively expensive. To bridge the gap, we investigate the weakly-supervised setting, where models learn from easily accessible video-language data without annotations. We identify that intra-sample spurious correlations among video-language components can be alleviated if the model captures the decomposed structures of video and language data. In this light, we propose a novel framework, namely WINNER, for hierarchical video-text understanding. WINNER first builds the language decomposition tree in a bottom-up manner, upon which the structural attention mechanism and top-down feature backtracking jointly build a multi-modal decomposition tree, permitting a hierarchical understanding of unstructured videos. The multi-modal decomposition tree serves as the basis for multi-hierarchy language-tube matching. A hierarchical contrastive learning objective is proposed to learn the multi-hierarchy correspondence and distinguishment with intra-sample and inter-sample video-text decomposition structures, achieving video-language decomposition structure alignment. Extensive experiments demonstrate the rationality of our design and its effectiveness beyond state-of-the-art weakly supervised methods, even some supervised methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94347882-bf07-4341-a9da-9f73c190af7fCited by top-tier papers23
- Panoptic Scene Graph Generation with Semantics-Prototype LearningLi Li, Wei Ji, Yiming Wu, Mengze Li et al.AAAI 2024 · 63 citations
- Intelligent Model Update Strategy for Sequential RecommendationZheqi Lv, Wenqiao Zhang, Zhengyu Chen, Shengyu Zhang et al.WWW 2024 · 53 citations
- Precedent-Enhanced Legal Judgment Prediction with LLM and Domain-Model CollaborationYiquan Wu, Siying Zhou, Yifei Liu, Weiming Lu et al.EMNLP 2023 · 37 citations
- Gradient-Regulated Meta-Prompt Learning for Generalizable Vision-Language ModelsJuncheng Li, Minghe Gao, Longhui Wei, Siliang Tang et al.ICCV 2023 · 34 citations
- Revisiting the Domain Shift and Sample Uncertainty in Multi-source Active Domain TransferWenqiao Zhang, Zheqi LvCVPR 2024 · 17 citations
Builds on33
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
- CauseRec: Counterfactual User Sequence Synthesis for Sequential RecommendationShengyu Zhang, Dong Yao, Zhou Zhao, Tat-Seng Chua et al.SIGIR 2021 · 118 citations
- BoostMIS: Boosting Medical Image Semi-supervised Learning with Adaptive Pseudo Labeling and Informative Active AnnotationWenqiao Zhang, Lei Zhu, James Hallinan, Shengyu Zhang et al.CVPR 2022 · 115 citations
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng et al.CVPR 2022 · 108 citations
Related papers
- TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video GroundingJinxuan Li, Yi Zhang, Jian-Fang Hu, Chaolei Tan et al.AAAI 2026 · 1 citation
- Video-Text Prompting for Weakly Supervised Spatio-Temporal Video GroundingHeng Zhao, Yinjie Zhao, Bihan Wen, Yew-Soon Ong et al.EMNLP 2024
- Learning Multi-Scale Video-Text Correspondence for Weakly Supervised Temporal Article GrondingWenjia Geng, Yong Liu, Lei Chen, Sujia Wang et al.AAAI 2024 · 3 citations
- Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph GroundingChaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang HuCVPR 2024 · 5 citations
- Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video GroundingYang Jin, Yongzhi Li, Zehuan Yuan, Yadong MuNeurIPS 2022 · 69 citations
