TF-Blender: Temporal Feature Blender for Video Object Detection
Yiming Cui, Liqi Yan, Zhiwen Cao, Dongfang Liu
Abstract
Video objection detection is a challenging task because isolated video frames may encounter appearance deterioration, which introduces great confusion for detection. One of the popular solutions is to exploit the temporal information and enhance per-frame representation through aggregating features from neighboring frames. Despite achieving improvements in detection, existing methods focus on the selection of higher-level video frames for aggregation rather than modeling lower-level temporal relations to increase the feature representation. To address this limitation, we propose a novel solution named TF-Blender, which includes three modules: 1) Temporal relation models the relations between the current frame and its neigh-boring frames to preserve spatial information. 2). Feature adjustment enriches the representation of every neigh-boring feature map; 3) Feature blender combines outputs from the first two modules and produces stronger features for the later detection tasks. For its simplicity, TF-Blender can be effortlessly plugged into any detection network to improve detection behavior. Extensive evaluations on ImageNet VID and YouTube-VIS benchmarks indicate the performance guarantees of using TF-Blender on recent state-of-the-art methods. Code is available at https://github.com/goodproj13/TF-Blender.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c47daa78-b4f1-4e69-aa48-9ec2366a2130Cited by top-tier papers17
- ClusterFomer: Clustering As A Universal Visual LearnerJames Liang, Yiming Cui, Qifan Wang, Tong Geng et al.NeurIPS 2023 · 63 citations
- Fusion Is Not Enough: Single Modal Attacks on Fusion Models for 3D Object DetectionZhiyuan Cheng, Hongjun Choi, Shiwei Feng, James Chenhao Liang et al.ICLR 2024 · 32 citations
- Explore Spatio-temporal Aggregation for Insubstantial Object Detection: Benchmark Dataset and BaselineKailai Zhou, Yibo Wang, Tao Lv, Yunqian Li et al.CVPR 2022 · 21 citations
- Provably Efficient Offline Reinforcement Learning for Partially Observable Markov Decision ProcessesHongyi Guo, Qi Cai, Yufeng Zhang, Zhuoran Yang et al.ICML 2022 · 17 citations
- MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target DetectionMingjin Zhang, Yuanjun Ouyang, Fei Gao, Jie Guo et al.AAAI 2025 · 10 citations
Builds on11
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Sequence Level Semantics Aggregation for Video Object DetectionHaiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 236 citations
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou et al.ICCV 2019 · 211 citations
- DenserNet: Weakly Supervised Visual Localization Using Multi-Scale Feature AggregationDongfang Liu, Yiming Cui, Liqi Yan, Christos Mousas et al.AAAI 2021 · 149 citations
Related papers
- Temporal Context Enhanced Feature Aggregation for Video Object DetectionFei He, Naiyu Gao, Qiaozhe Li, Senyao Du et al.AAAI 2020 · 40 citations
- TGBFormer: Transformer-GraphFormer Blender Network for Video Object DetectionQiang Qi, Xiao WangAAAI 2025 · 5 citations
- Feature Aggregated Queries for Transformer-Based Video Object DetectorsYiming CuiCVPR 2023
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
- Exploiting Better Feature Aggregation for Video Object DetectionLiang Han, Pichao Wang, Zhaozheng Yin, Fan Wang et al.ACM MM 2020 · 37 citations
