Online Model Distillation for Efficient Video Inference
Ravi Teja Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan, Kayvon Fatahalian
摘要
High-quality computer vision models typically address the problem of understanding the general distribution of real-world images. However, most cameras observe only a very small fraction of this distribution. This offers the possibility of achieving more efficient inference by specializing compact, low-cost models to the specific distribution of frames observed by a single camera. In this paper, we employ the technique of model distillation (supervising a low-cost student model using the output of a high-cost teacher) to specialize accurate, low-cost semantic segmentation models to a target video stream. Rather than learn a specialized student model on offline data from the video stream, we train the student in an online fashion on the live video, intermittently running the teacher to provide a target for learning. Online model distillation yields semantic segmentation models that closely approximate their Mask R-CNN teacher with 7 to 17 lower inference runtime cost (11 to 26 in FLOPs), even when the target video's distribution is non-stationary. Our method requires no offline pretraining on the target video stream, achieves higher accuracy and lower cost than solutions based on flow or video object segmentation, and can exhibit better temporal stability than the original teacher. We also provide a new video dataset for evaluating the efficiency of inference over long running video streams.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
- Titans: Learning to Memorize at Test TimeAli Behrouz, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 被引用 368 次
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 被引用 356 次
- Test-Time Training with Masked AutoencodersYossi Gandelsman, Yu Sun, Xinlei Chen, Alexei A. EfrosNeurIPS 2022 · 被引用 283 次
- Hand-Object Contact Consistency Reasoning for Human Grasps GenerationHanwen Jiang, Shaowei Liu, Jiashun Wang, Xiaolong WangICCV 2021 · 被引用 242 次
相关 Paper
- Real-Time Video Inference on Edge Devices via Adaptive Model StreamingMehrdad Khani Shirkoohi, Pouya Hamadanian, Arash Nasr-Esfahany, Mohammad AlizadehICCV 2021 · 被引用 58 次
- WeClick: Weakly-Supervised Video Semantic Segmentation with Click AnnotationsPeidong Liu, Zibin He, Xiyu Yan, Yong Jiang 等ACM MM 2021 · 被引用 9 次
- How many Observations are Enough? Knowledge Distillation for Trajectory ForecastingAlessio Monti, Angelo Porrello, Simone Calderara, Pasquale Coscia 等CVPR 2022 · 被引用 62 次
- Lightweight Spatio-Temporal Modeling via Temporally Shifted Distillation for Real-Time Accident AnticipationPatrik Patera, Yie-Tarng Chen, Wen-Hsien FangICLR 2026
- DistInit: Learning Video Representations Without a Single Labeled VideoRohit Girdhar, Du Tran, Lorenzo Torresani, Deva RamananICCV 2019 · 被引用 59 次
