CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding
Zhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao, Kun Yan, Wing Kwong Chan, Chong-Wah Ngo, Mike Zheng Shou, Nan Duan
摘要
This paper tackles an emerging and challenging problem of long video temporal grounding (VTG) that localizes video moments related to a natural language (NL) query. Compared with short videos, long videos are also highlydemanded but less explored, which brings new challenges in higher inference computation cost and weaker multi-modal alignment. To address these challenges, we propose CONE, an efficient COarse-to-fiNE alignment framework. CONE is a plug-and-play framework on top of existing VTG models to handle long videos through a sliding window mechanism. Specifically, CONE (1) introduces a query-guided window selection strategy to speed up inference, and (2) proposes a coarse-to-fine mechanism via a novel incorporation of contrastive learning to enhance multi-modal alignment for long videos. Extensive experiments on two large-scale long VTG benchmarks consistently show both substantial performance gains (e.g., 3.13% +119% ---→6.87% on MAD) and state-of-theart results. Analyses also reveal higher efficiency as the query-guided window selection mechanism accelerates inference time by 2x on Ego4D-NLQ and 15x on MAD while keeping SOTA results. Codes have been released at https://github.com/houzhijian/CONE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang 等NeurIPS 2025 · 被引用 30 次
- Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long VideosYulin Pan, Xiangteng He, Biao Gong, Yiliang Lv 等ICCV 2023 · 被引用 29 次
- Boundary Denoising for Video Activity LocalizationMengmeng Xu, Mattia Soldan, Jialin Gao, Shuming Liu 等ICLR 2024 · 被引用 16 次
- SnAG: Scalable and Accurate Video GroundingFangzhou Mu, Sicheng Mo, Yin LiCVPR 2024 · 被引用 13 次
- Video-Context Aligned Transformer for Video Question AnsweringLinlin Zong, Jiahui Wan, Xianchao Zhang, Xinyue Liu 等AAAI 2024 · 被引用 8 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 被引用 425 次
相关 Paper
- Localizing Moments in Long Video Via Multimodal GuidanceWayner Barrios, Mattia Soldan, Alberto Mario Ceballos-Arroyo, Fabian Caba Heilbron 等ICCV 2023 · 被引用 32 次
- Multi-Scale Contrastive Learning for Video Temporal GroundingThong Thanh Nguyen, Yi Bin, Xiaobao Wu, Zhiyuan Hu 等AAAI 2025 · 被引用 7 次
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng 等CVPR 2022 · 被引用 108 次
- CausalVTG: Towards Robust Video Temporal Grounding via Causal InferenceQiyi Wang, Senda Chen, Ying ShenNeurIPS 2025 · 被引用 1 次
- Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal GroundingMinseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim 等NeurIPS 2025 · 被引用 4 次
