Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding
Chaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang Hu
摘要
Video Paragraph Grounding (VPG) is an emerging task in video-language understanding, which aims at localizing multiple sentences with semantic relations and temporal or-der from an untrimmed video. However, existing VPG approaches are heavily reliant on a considerable number of temporal labels that are laborious and time-consuming to acquire. In this work, we introduce and explore Weakly-Supervised Video Paragraph Grounding (WSVPG) to elim-inate the need of temporal annotations. Different from pre-vious weakly-supervised grounding frameworks based on multiple instance learning or reconstruction learning for two-stage candidate ranking, we propose a novel siamese learning framework that jointly learns the cross-modal feature alignment and temporal coordinate regression without timestamp labels to achieve concise one-stage localization for WSVPG. Specifically, we devise a Siamese Grounding TRansformer (SiamGTR) consisting of two weight-sharing branches for learning complementary supervision. An Aug-mentation Branch is utilized for directly regressing the tem-poral boundaries of a complete paragraph within a pseudo video, and an Inference Branch is designed to capture the order-guided feature correspondence for localizing multi-ple sentences in a normal video. We demonstrate by exten-sive experiments that our paradigm has superior practica-bility and flexibility to achieve efficient weakly-supervised or semi-supervised learning, outperforming state-of-the-art methods trained with the same or stronger supervision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Enhancing Partially Relevant Video Retrieval with Hyperbolic LearningJun Li, Jinpeng Wang, Chaolei Tan, Niu Lian 等ICCV 2025 · 被引用 5 次
- Enrich and Detect: Video Temporal Grounding With Multimodal LlmsShraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa 等ICCV 2025 · 被引用 4 次
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi 等ACM MM 2024 · 被引用 2 次
- Object-Shot Enhanced Grounding Network for Egocentric VideoYisen Feng, Haoyu Zhang, Meng Liu, Weili Guan 等CVPR 2025
- Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video GroundingXin Gu, Yaojie Shen, Chenxi Luo, Tiejian Luo 等ICLR 2025
它引用的顶会 Paper55
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
相关 Paper
- Semi-supervised Video Paragraph Grounding with Contrastive EncoderXun Jiang, Xing Xu, Jingran Zhang, Fumin Shen 等CVPR 2022 · 被引用 42 次
- Learning Multi-Scale Video-Text Correspondence for Weakly Supervised Temporal Article GrondingWenjia Geng, Yong Liu, Lei Chen, Sujia Wang 等AAAI 2024 · 被引用 3 次
- Cross-Sentence Temporal and Semantic Relations in Video Activity LocalisationJiabo Huang, Yang Liu, Shaogang Gong, Hailin JinICCV 2021 · 被引用 77 次
- STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingRui Su, Qian Yu, Dong XuICCV 2021 · 被引用 75 次
- WINNER: Weakly-supervised hIerarchical decompositioN and aligNment for spatio-tEmporal video gRoundingMengze Li, Han Wang, Wenqiao Zhang, Jiaxu Miao 等CVPR 2023
