System-Status-Aware Adaptive Network for Online Streaming Video Understanding
Lin Geng Foo, Jia Gong, Zhipeng Fan, Jun Liu
摘要
Recent years have witnessed great progress in deep neural networks for real-time applications. However, most existing works do not explicitly consider the general case where the device's state and the available resources fluctuate over time, and none of them investigate or address the impact of varying computational resources for online video understanding tasks. This paper proposes a System-status-aware Adaptive Network (SAN) that considers the device's real-time state to provide high-quality predictions with low delay. Usage of our agent's policy improves efficiency and robustness to fluctuations of the system status. On two widely used video understanding tasks, SAN obtains state-of-the-art performance while constantly keeping processing delays low. Moreover, training such an agent on various types of hardware configurations is not easy as the labeled training data might not be available, or can be computationally prohibitive. To address this challenging problem, we propose a Meta Self-supervised Adaptation (MSA) method that adapts the agent's policy to new hardware configurations at test-time, allowing for easy deployment of the model onto other unseen hardware platforms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Generative Model-Based Feature Knowledge Distillation for Action RecognitionGuiqin Wang, Peng Zhao, Yanjiang Shi, Cong Zhao 等AAAI 2024 · 被引用 9 次
- : Discrete Diffusion Model for Occluded 3D Human Pose EstimationWeiquan Wang, Jun Xiao, Chunping Wang, Wei Liu 等NeurIPS 2024 · 被引用 4 次
- Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal GroundingZelin Zheng, Xinyan Liu, Ruixin Li, Antoni B. Chan 等ICML 2026 · 被引用 1 次
- TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight DetectionHao Sun, Mingyao Zhou, Wenjing Chen, Wei XieAAAI 2024
- Number it: Temporal Grounding Videos like Flipping MangaYongliang Wu, Xinting Hu, Yuyang Sun, Yizhou Zhou 等CVPR 2025
它引用的顶会 Paper12
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 被引用 4,239 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai 等ACL 2020 · 被引用 215 次
相关 Paper
- Scene-Adaptive Video Frame Interpolation via Meta-LearningMyungsub Choi, Janghoon Choi, Sungyong Baik, Tae Hyun Kim 等CVPR 2020
- CASVA: Configuration-Adaptive Streaming for Live Video AnalyticsMiao Zhang, Fangxin Wang, Jiangchuan LiuINFOCOM 2022 · 被引用 69 次
- DeDelayed: Deleting Remote Inference Delay via On-Device CorrectionDan Jacobellis, Mateen Ulhaq, Fabien Racapé, Hyomin Choi 等CVPR 2026 · 被引用 2 次
- Hardware-adaptive Efficient Latency Prediction for NAS via Meta-LearningHayeon Lee, Sewoong Lee, Song Chong, Sung Ju HwangNeurIPS 2021 · 被引用 32 次
- PASS: Patch Automatic Skip Scheme for Efficient Real-Time Video Perception on Edge DevicesQihua Zhou, Song Guo, Jun Pan, Jiacheng Liang 等AAAI 2023
