Streaming Video Model
Yucheng Zhao, Chong Luo, Chuanxin Tang, Dongdong Chen, Noel Codella, Zheng-Jun Zha
摘要
Figure 1. Illustration of the proposed streaming video model with a comparison to conventional frame-based architecture and clip-based architecture. (a) The two-stage streaming video model gracefully serves different types of video tasks through a unified architecture. The output of the temporal-aware (T-aware) spatial encoder serves the frame-based tasks, such as MOT, while the output of the temporal decoder serves the sequence-based tasks, such as action recognition. (b) Frame-based architecture, which uses single image model to independently extract spatial features for each frame, is widely used in the frame-based video tasks. (c) Clip-based architecture, which uses video model to produce the spatiotemporal features for an entire clip, is widely used in the sequence-based video tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Streaming Dense Video CaptioningXingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan 等CVPR 2024 · 被引用 33 次
- Streaming Video Instruction TuningJiaer Xia, Peixian Chen, Mengdan Zhang, Xing Sun 等CVPR 2026 · 被引用 28 次
- A Multimodal, Multi-Task Adapting Framework for Video Action RecognitionMengmeng Wang, Jiazheng Xing, Boyuan Jiang, Jun Chen 等AAAI 2024 · 被引用 13 次
- Action Detail Matters: Refining Video Recognition with Local Action QueriesMengmeng Wang, Zeyi Huang, Xiangjie Kong, Guojiang Shen 等CVPR 2025
- LiveCC: Learning Video LLM with Streaming Speech Transcription at ScaleJoya Chen, Ziyun Zeng, Yiqi Lin, Wei Li 等CVPR 2025
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
相关 Paper
- A Deeper Dive Into What Deep Spatiotemporal Networks Encode: Quantifying Static vs. Dynamic InformationMatthew Kowal, Mennatullah Siam, Md. Amirul Islam, Neil D. B. Bruce 等CVPR 2022 · 被引用 23 次
- CAST: Cross-Attention in Space and Time for Video Action RecognitionDongho Lee, Jongseo Lee, Jinwoo ChoiNeurIPS 2023 · 被引用 43 次
- MAU: A Motion-Aware Unit for Video Prediction and BeyondZheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma 等NeurIPS 2021 · 被引用 193 次
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu 等ICCV 2019 · 被引用 442 次
- MOSO: Decomposing MOtion, Scene and Object for Video PredictionMingzhen Sun, Weining Wang, Xinxin Zhu, Jing LiuCVPR 2023
