Rethinking the Bottom-Up Framework for Query-Based Video Localization
Long Chen, Chujie Lu, Siliang Tang, Jun Xiao, Dong Zhang, Chilie Tan, Xiaolin Li
摘要
In this paper, we focus on the task query-based video localization, i.e., localizing a query in a long and untrimmed video. The prevailing solutions for this problem can be grouped into two categories: i) Top-down approach: It pre-cuts the video into a set of moment candidates, then it does classification and regression for each candidate; ii) Bottom-up approach: It injects the whole query content into each video frame, then it predicts the probabilities of each frame as a ground truth segment boundary (i.e., start or end). Both two frameworks have respective shortcomings: the top-down models suffer from heavy computations and they are sensitive to the heuristic rules, while the performance of bottom-up models is behind the performance of top-down counterpart thus far. However, we argue that the performance of bottom-up framework is severely underestimated by current unreasonable designs, including both the backbone and head network. To this end, we design a novel bottom-up model: Graph-FPN with Dense Predictions (GDP). For the backbone, GDP firstly generates a frame feature pyramid to capture multi-level semantics, then it utilizes graph convolution to encode the plentiful scene relationships, which incidentally mitigates the semantic gaps in the multi-scale feature pyramid. For the head network, GDP regards all frames falling in the ground truth segment as the foreground, and each foreground frame regresses the unique distances from its location to bi-directional boundaries. Extensive experiments on two challenging query-based video localization tasks (natural language video localization and video relocalization), involving four challenging benchmarks (TACoS, Charades-STA, ActivityNet Captions, and Activity-VRL), have shown that GDP surpasses the state-of-the-art top-down models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- Boundary Proposal Network for Two-stage Natural Language Video LocalizationShaoning Xiao, Long Chen, Songyang Zhang, Wei Ji 等AAAI 2021 · 被引用 186 次
- Negative Sample Matters: A Renaissance of Metric Learning for Temporal GroundingZhenzhi Wang, Limin Wang, Tao Wu, Tianhao Li 等AAAI 2022 · 被引用 170 次
- Proposal-Free Video Grounding with Contextual Pyramid NetworkKun Li, Dan Guo, Meng WangAAAI 2021 · 被引用 138 次
- Fast Video Moment RetrievalJunyu Gao, Changsheng XuICCV 2021 · 被引用 132 次
- Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression GroundingLong Chen, Wenbo Ma, Jun Xiao, Hanwang Zhang 等AAAI 2021 · 被引用 118 次
它引用的顶会 Paper2
相关 Paper
- Dual Path Interaction Network for Video Moment LocalizationHao Wang, Zheng-Jun Zha, Xuejin Chen, Zhiwei Xiong 等ACM MM 2020 · 被引用 69 次
- Cascaded Prediction Network via Segment Tree for Temporal Video GroundingYang Zhao, Zhou Zhao, Zhu Zhang, Zhijie LinCVPR 2021
- Dense Regression Network for Video GroundingRunhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen 等CVPR 2020
- Adaptive Proposal Generation Network for Temporal Sentence Localization in VideosDaizong Liu, Xiaoye Qu, Jianfeng Dong, Pan ZhouEMNLP 2021 · 被引用 40 次
- Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 等CVPR 2021
