DeCo: Decomposition and Reconstruction for Compositional Temporal Grounding via Coarse-to-Fine Contrastive Ranking
Lijin Yang, Quan Kong, Hsuan-Kung Yang, Wadim Kehl, Yoichi Sato, Norimasa Kobori
Abstract
Understanding dense action in videos is a fundamental challenge towards the generalization of vision models. Several works show that compositionality is key to achieving generalization by combining known primitive elements, especially for handling novel composited structures. Compositional temporal grounding is the task of localizing dense action by using known words combined in novel ways in the form of novel query sentences for the actual grounding. In recent works, composition is assumed to be learned from pairs of whole videos and language embeddings through large scale self-supervised pre-training. Alternatively, one can process the video and language into word-level primitive elements, and then only learn fine-grained semantic correspondences. Both approaches do not consider the granularity of the compositions, where different query granularity corresponds to different video segments. Therefore, a good compositional representation should be sensitive to different video and query granularity. We propose a method to learn a coarse-to-fine compositional representation by decomposing the original query sentence into different granular levels, and then learning the correct correspondences between the video and recombined queries through a contrastive ranking constraint. Additionally, we run temporal boundary prediction in a coarse-to-fine manner for precise grounding boundary detection. Experiments are performed on two datasets, Charades-CG and ActivityNet-CG, showing the superior compositional generalizability of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a81e710d-9857-4980-b33a-9119ccd411cfCited by top-tier papers8
- PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkMingyao Zhou, Hao Sun, Wei Xie, Ming Dong et al.NeurIPS 2025 · 1 citation
- In-Context Compositional Generalization for Large Vision-Language ModelsChuanhao Li, Chenchen Jing, Zhen Li, Mingliang Zhai et al.EMNLP 2024 · 1 citation
- Consistency of Compositional Generalization Across Multiple LevelsChuanhao Li, Zhen Li, Chenchen Jing, Xiaomeng Fan et al.AAAI 2025 · 1 citation
- STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal GroundingAaryan Garg, Akash Kumar, Yogesh S. RawatCVPR 2025
- Localizing Events in Videos with Multimodal QueriesGengyuan Zhang, Mang Ling Ada Fok, Jialu Ma, Yan Xia et al.CVPR 2025
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 279 citations
- Boundary Proposal Network for Two-stage Natural Language Video LocalizationShaoning Xiao, Long Chen, Songyang Zhang, Wei Ji et al.AAAI 2021 · 186 citations
Related papers
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence GroundingJiaming Chen, Weixin Luo, Wei Zhang, Lin MaAAAI 2022 · 33 citations
- Tree-Structured Policy Based Progressive Reinforcement Learning for Temporally Language Grounding in VideoJie Wu, Guanbin Li, Si Liu, Liang LinAAAI 2020 · 117 citations
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng et al.CVPR 2022 · 108 citations
- Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed VideosJie Wu, Guanbin Li, Xiaoguang Han, Liang LinACM MM 2020 · 70 citations
