Weakly-Supervised Temporal Article Grounding
Long Chen, Yulei Niu, Brian Chen, Xudong Lin, Guangxing Han, Christopher Thomas, Hammad A. Ayyubi, Heng Ji, Shih-Fu Chang
Abstract
Given a long untrimmed video and natural language queries, video grounding (VG) aims to temporally localize the semantically-aligned video segments. Almost all existing VG work holds two simple but unrealistic assumptions: 1) All query sentences can be grounded in the corresponding video. 2) All query sentences for the same video are always at the same semantic scale. Unfortunately, both assumptions make today's VG models fail to work in practice. For example, in real-world multimodal assets (e.g., news articles), most of the sentences in the article can not be grounded in their affiliated videos, and they typically have rich hierarchical relations (i.e., at different semantic scales). To this end, we propose a new challenging grounding task: Weakly-Supervised temporal Article Grounding (WSAG). Specifically, given an article and a relevant video, WSAG aims to localize all "groundable" sentences to the video, and these sentences are possibly at different semantic scales. Accordingly, we collect the first WSAG dataset to facilitate this task: Youwiki-How, which borrows the inherent multi-scale descriptions in wikiHow articles and plentiful YouTube videos. In addition, we propose a simple but effective method DualMIL for WSAG, which consists of a two-level MIL 1 loss and a single-/cross-sentence constraint loss. These training objectives are carefully designed for these relaxed assumptions. Extensive ablations have verified the effectiveness of DualMIL 2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 755d4706-150a-4a1f-ab09-283087430268Cited by top-tier papers7
- Learning to Ground Instructional Articles in Videos through NarrationsEffrosyni Mavroudi, Triantafyllos Afouras, Lorenzo TorresaniICCV 2023 · 28 citations
- Non-Sequential Graph Script Induction via Multimedia GroundingYu Zhou, Sha Li, Manling Li, Xudong Lin et al.ACL 2023 · 5 citations
- Learning Multi-Scale Video-Text Correspondence for Weakly Supervised Temporal Article GrondingWenjia Geng, Yong Liu, Lei Chen, Sujia Wang et al.AAAI 2024 · 3 citations
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi et al.ACM MM 2024 · 2 citations
- STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal GroundingAaryan Garg, Akash Kumar, Yogesh S. RawatCVPR 2025
Builds on13
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 279 citations
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 206 citations
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang et al.AAAI 2020 · 170 citations
Related papers
- Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph GroundingChaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang HuCVPR 2024 · 5 citations
- Activity-driven Weakly-Supervised Spatio-Temporal Grounding from Untrimmed VideosJunwen Chen, Wentao Bao, Yu KongACM MM 2020 · 19 citations
- Unsupervised Temporal Video Grounding with Deep Semantic ClusteringDaizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di et al.AAAI 2022 · 52 citations
- D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance AnnotationHanjun Li, Xiujun Shu, Sunan He, Ruizhi Qiao et al.ICCV 2023 · 21 citations
- Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal GroundingMinseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim et al.NeurIPS 2025 · 4 citations
