Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning
Juncheng Li, Junlin Xie, Long Qian, Linchao Zhu, Siliang Tang, Fei Wu, Yi Yang, Yueting Zhuang, Xin Eric Wang
Abstract
Temporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language descriptions, temporal grounding allows activity grounding beyond pre-defined classes and has received increasing attention in recent years. The semantic diversity is rooted in the principle of compositionality in linguistics, where novel semantics can be systematically described by combining known words in novel ways (compositional generalization). However, current temporal grounding datasets do not specifically test for the compositional generalizability. To systematically measure the compositional generalizability of temporal grounding models, we introduce a new Compositional Temporal Grounding task and construct two new dataset splits, i.e., Charades-CG and ActivityNet-CG. Evaluating the state-of-the-art methods on our new dataset splits, we empirically find that they fail to generalize to queries with novel combinations of seen words. To tackle this challenge, we propose a variational cross-graph reasoning framework that explicitly decomposes video and language into multiple structured hierarchies and learns fine-grained semantic correspondence among them. Experiments illustrate the superior compositional generalizability of our approach. The repository of this work is at ht tps: / / gi thub. com/YYJMJC/ Composi tional- Temporal-Grounding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46e80558-c0cc-4c03-8e62-07a313a5a1f6Cited by top-tier papers37
- Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningLong Qian, Juncheng Li, Yu Wu, Yaobo Ye et al.ICML 2024 · 121 citations
- BoostMIS: Boosting Medical Image Semi-supervised Learning with Adaptive Pseudo Labeling and Informative Active AnnotationWenqiao Zhang, Lei Zhu, James Hallinan, Shengyu Zhang et al.CVPR 2022 · 115 citations
- Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsJuncheng Li, Kaihang Pan, Zhiqi Ge, Minghe Gao et al.ICLR 2024 · 95 citations
- Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open WorldQifan Yu, Juncheng Li, Yu Wu, Siliang Tang et al.ICCV 2023 · 51 citations
- End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACL 2022 · 46 citations
Builds on22
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 279 citations
- Fine-Grained Action Retrieval Through Multiple Parts-of-Speech EmbeddingsMichael Wray, Gabriela Csurka, Diane Larlus, Dima DamenICCV 2019 · 185 citations
- Learning Compositional Rules via Neural Program SynthesisMaxwell I. Nye, Armando Solar-Lezama, Josh Tenenbaum, Brenden M. LakeNeurIPS 2020 · 120 citations
- CauseRec: Counterfactual User Sequence Synthesis for Sequential RecommendationShengyu Zhang, Dong Yao, Zhou Zhao, Tat-Seng Chua et al.SIGIR 2021 · 118 citations
Related papers
- DeCo: Decomposition and Reconstruction for Compositional Temporal Grounding via Coarse-to-Fine Contrastive RankingLijin Yang, Quan Kong, Hsuan-Kung Yang, Wadim Kehl et al.CVPR 2023
- Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence GroundingJiaming Chen, Weixin Luo, Wei Zhang, Lin MaAAAI 2022 · 33 citations
- HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in VideosTingting Han, Xinsong Tao, Yufei Yin, Min Tan et al.CVPR 2026
- Embracing Uncertainty: Decoupling and De-Bias for Robust Temporal GroundingHao Zhou, Chongyang Zhang, Yan Luo, Yanjun Chen et al.CVPR 2021
- Consistency of Compositional Generalization Across Multiple LevelsChuanhao Li, Zhen Li, Chenchen Jing, Xiaomeng Fan et al.AAAI 2025 · 1 citation
