Streaming Video Model
Yucheng Zhao, Chong Luo, Chuanxin Tang, Dongdong Chen, Noel Codella, Zheng-Jun Zha
Abstract
Figure 1. Illustration of the proposed streaming video model with a comparison to conventional frame-based architecture and clip-based architecture. (a) The two-stage streaming video model gracefully serves different types of video tasks through a unified architecture. The output of the temporal-aware (T-aware) spatial encoder serves the frame-based tasks, such as MOT, while the output of the temporal decoder serves the sequence-based tasks, such as action recognition. (b) Frame-based architecture, which uses single image model to independently extract spatial features for each frame, is widely used in the frame-based video tasks. (c) Clip-based architecture, which uses video model to produce the spatiotemporal features for an entire clip, is widely used in the sequence-based video tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Streaming Dense Video CaptioningXingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan et al.CVPR 2024 · 33 citations
- Streaming Video Instruction TuningJiaer Xia, Peixian Chen, Mengdan Zhang, Xing Sun et al.CVPR 2026 · 28 citations
- A Multimodal, Multi-Task Adapting Framework for Video Action RecognitionMengmeng Wang, Jiazheng Xing, Boyuan Jiang, Jun Chen et al.AAAI 2024 · 13 citations
- Action Detail Matters: Refining Video Recognition with Local Action QueriesMengmeng Wang, Zeyi Huang, Xiangjie Kong, Guojiang Shen et al.CVPR 2025
- LiveCC: Learning Video LLM with Streaming Speech Transcription at ScaleJoya Chen, Ziyun Zeng, Yiqi Lin, Wei Li et al.CVPR 2025
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
Related papers
- A Deeper Dive Into What Deep Spatiotemporal Networks Encode: Quantifying Static vs. Dynamic InformationMatthew Kowal, Mennatullah Siam, Md. Amirul Islam, Neil D. B. Bruce et al.CVPR 2022 · 23 citations
- CAST: Cross-Attention in Space and Time for Video Action RecognitionDongho Lee, Jongseo Lee, Jinwoo ChoiNeurIPS 2023 · 43 citations
- MAU: A Motion-Aware Unit for Video Prediction and BeyondZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma et al.NeurIPS 2021 · 193 citations
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
- MOSO: Decomposing MOtion, Scene and Object for Video PredictionMingzhen Sun, Weining Wang, Xinxin Zhu, Jing LiuCVPR 2023
