Frame Flexible Network
Yitian Zhang, Yue Bai, Chang Liu, Huan Wang, Sheng Li, Yun Fu
Abstract
Existing video recognition algorithms always conduct different training pipelines for inputs with different frame numbers, which requires repetitive training operations and multiplying storage costs. If we evaluate the model using other frames which are not used in training, we observe the performance will drop significantly (see Fig. 1 ), which is summarized as Temporal Frequency Deviation phenomenon. To fix this issue, we propose a general framework, named Frame Flexible Network (FFN), which not only enables the model to be evaluated at different frames to adjust its computation, but also reduces the memory costs of storing multiple models significantly. Concretely, FFN integrates several sets of training sequences, involves Multi-Frequency Alignment (MFAL) to learn temporal frequency invariant representations, and leverages Multi-Frequency Adaptation (MFAD) to further strengthen the representation abilities. Comprehensive empirical validations using various architectures and popular benchmarks solidly demonstrate the effectiveness and generalization of FFN (e.g., 7.08/5.15/2.17% performance gain at Frame 4/8/16 on Something-Something V1 dataset over Uniformer). Code is available at https://github.com/BeSpontaneous/FFN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- TDN: Temporal Difference Networks for Efficient Action RecognitionLimin Wang, Zhan Tong, Bin Ji, Gangshan WuCVPR 2021
- Frame2Freq: Spectral Adapters for Fine-Grained Video UnderstandingThinesh Thiyakesan Ponbagavathi, Constantin Seibold, Alina RoitbergCVPR 2026 · 2 citations
- FASTER Recurrent Networks for Efficient Video ClassificationLinchao Zhu, Du Tran, Laura Sevilla-Lara, Yi Yang et al.AAAI 2020 · 60 citations
- State Space Models for Event CamerasNikola Zubic, Mathias Gehrig, Davide ScaramuzzaCVPR 2024 · 33 citations
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation NetworkQiang Qi, Wenqi Shang, Meifang Wang, Xiao WangCVPR 2026
