Weakly-Supervised Video Object Grounding via Stable Context Learning
Wei Wang, Junyu Gao, Changsheng Xu
Abstract
We investigate the problem of weakly-supervised video object grounding (WSVOG), where only the video-sentence annotations are provided for training. It aims at localizing the queried objects described in the sentence to visual regions in the video. Despite the recent progress, existing approaches have not fully exploited the potential of the description sentences for cross-modal alignment in two aspects: (1) Most of them extract objects from the description sentences and represent them with fixed textual representations. While achieving promising results, they do not make full use of the contextual information in the sentence. (2) A few works have attempted to utilize contextual information to learn object representations, but found a significant decrease in performance due to the unstable training in cross-modal alignment. To address the above issues, in this paper, we propose a Stable Context Learning (SCL) framework for WSVOG which jointly enjoys the merits of stable learning and rich contextual information. Specifically, we design two modules named Context-Aware Object Stabilizer module and Cross-Modal Alignment Knowledge Transfer module, which are cooperated together to inject contextual information to stable object concepts in text modality and transfer contextualized knowledge in cross-modal alignment. Our approach is finally optimized under a frame-level MIL paradigm. Extensive experiments on three popular benchmarks demonstrate its significant effectiveness.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a1271281-66e2-42a7-bb4a-beac04119489Cited by top-tier papers5
- HERO: HiErarchical spatio-tempoRal reasOning with Contrastive Action Correspondence for End-to-End Video Object GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACM MM 2022 · 25 citations
- Where Does It Exist from the Low-Altitude: Spatial Aerial Video GroundingYang Zhan, Yuan YuanNeurIPS 2025 · 8 citations
- Weakly-Supervised Generation and Grounding of Visual Descriptions with Conditional Generative ModelsEffrosyni Mavroudi, René VidalCVPR 2022 · 5 citations
- Inverse Compositional Learning for Weakly-supervised Relation GroundingHuan Li, Ping Wei, Zeyu Ma, Nanning ZhengICCV 2023 · 1 citation
- Learning to Segment Referred Objects from Narrated Egocentric VideosYuhan Shen, Huiyu Wang, Xitong Yang, Matt Feiszli et al.CVPR 2024 · 1 citation
Related papers
- Learning Multi-Scale Video-Text Correspondence for Weakly Supervised Temporal Article GrondingWenjia Geng, Yong Liu, Lei Chen, Sujia Wang et al.AAAI 2024 · 3 citations
- Weakly-supervised Video Scene Graph Generation via Unbiased Cross-modal LearningZiyue Wu, Junyu Gao, Changsheng XuACM MM 2023 · 5 citations
- Visual Co-Occurrence Alignment Learning for Weakly-Supervised Video Moment RetrievalZheng Wang, Jingjing Chen, Yu-Gang JiangACM MM 2021 · 74 citations
- Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text UnderstandingJongbhin Woo, Hyeonggon Ryu, Youngjoon Jang, Jae-Won Cho et al.ACM MM 2024 · 3 citations
- AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual GroundingYidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao et al.ACM MM 2025 · 2 citations
