You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed Videos
Xiang Fang, Daizong Liu, Pan Zhou, Guoshun Nan
Abstract
Given an untrimmed video, temporal sentence grounding (TSG) aims to locate a target moment semantically according to a sentence query. Although previous respectable works have made decent success, they only focus on highlevel visual features extracted from the consecutive decoded frames and fail to handle the compressed videos for query modelling, suffering from insufficient representation capability and significant computational complexity during training and testing. In this paper, we pose a new setting, compressed-domain TSG, which directly utilizes compressed videos rather than fully-decompressed frames as the visual input. To handle the raw video bit-stream input, we propose a novel Three-branch Compressed-domain Spatial-temporal Fusion (TCSF) framework, which extracts and aggregates three kinds of low-level visual features (Iframe, motion vector and residual features) for effective and efficient grounding. Particularly, instead of encoding the whole decoded frames like previous works, we capture the appearance representation by only learning the I-frame feature to reduce delay or latency. Besides, we explore the motion information not only by learning the motion vector feature, but also by exploring the relations of neighboring frames via the residual feature. In this way, a three-branch spatial-temporal attention layer with an adaptive motionappearance fusion module is further designed to extract and aggregate both appearance and motion information for the final grounding. Experiments on three challenging datasets shows that our TCSF achieves better performance than other state-of-the-art methods with lower complexity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers33
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang et al.ACM MM 2022 · 65 citations
- Panoptic Scene Graph Generation with Semantics-Prototype LearningLi Li, Wei Ji, Yiming Wu, Mengze Li et al.AAAI 2024 · 63 citations
- Pandora's Box: Towards Building Universal Attackers against Real-World Large Vision-Language ModelsDaizong Liu, Mingyu Yang, Xiaoye Qu, Pan Zhou et al.NeurIPS 2024 · 51 citations
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou et al.AAAI 2024 · 30 citations
- Combating Multimodal LLM Hallucination via Bottom-Up Holistic ReasoningShengqiong Wu, Hao Fei, Liangming Pan, William Yang Wang et al.AAAI 2025 · 24 citations
Builds on39
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 279 citations
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao et al.AAAI 2020 · 182 citations
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang et al.AAAI 2020 · 170 citations
Related papers
- Temporal Sentence Grounding in Streaming VideosTian Gan, Xiao Wang, Yan Sun, Jianlong Wu et al.ACM MM 2023 · 5 citations
- Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou et al.CVPR 2021
- Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence GroundingJiaming Chen, Weixin Luo, Wei Zhang, Lin MaAAAI 2022 · 33 citations
- Unsupervised Temporal Video Grounding with Deep Semantic ClusteringDaizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di et al.AAAI 2022 · 52 citations
- D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance AnnotationHanjun Li, Xiujun Shu, Sunan He, Ruizhi Qiao et al.ICCV 2023 · 21 citations
