MobileVOS: Real-Time Video Object Segmentation Contrastive Learning meets Knowledge Distillation
Roy Miles, Mehmet Kerim Yucel, Bruno Manganelli, Albert Saà-Garriga
Abstract
This paper tackles the problem of semi-supervised video object segmentation on resource-constrained devices, such as mobile phones. We formulate this problem as a distillation task, whereby we demonstrate that small spacetime-memory networks with finite memory can achieve competitive results with state of the art, but at a fraction of the computational cost (32 milliseconds per frame on a Samsung Galaxy S22). Specifically, we provide a theoretically grounded framework that unifies knowledge distillation with supervised contrastive representation learning. These models are able to jointly benefit from both pixel-wise contrastive learning and distillation from a pretrained teacher. We validate this loss by achieving competitive J &F to state of the art on both the standard DAVIS and YouTube benchmarks, despite running up to ˆ5 faster, and with ˆ32 fewer parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Understanding the Role of the Projector in Knowledge DistillationRoy Miles, Krystian MikolajczykAAAI 2024 · 60 citations
- DACAPO: Accelerating Continuous Learning in Autonomous Systems for Video AnalyticsYoonsung Kim, Changhun Oh, Jinwoo Hwang, Wonung Kim et al.ISCA 2024 · 13 citations
- : Improving Knowledge Distillation Using Orthogonal ProjectionsRoy Miles, Ismail Elezi, Jiankang DengCVPR 2024 · 9 citations
- InfoSAM: Fine-Tuning the Segment Anything Model from An Information-Theoretic PerspectiveYuanhong Zhang, Muyao Yuan, Weizhan Zhang, Tieliang Gong et al.ICML 2025
- Procedure Knowledge Decoupled Distillation Strategy for Procedure Planning in Instructional VideosXiaotian Pan, Zhaobo Qi, Xin Sun, Yuanrong Xu et al.AAAI 2025
Builds on28
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
Related papers
- Learning Position and Target Consistency for Memory-Based Video Object SegmentationLi Hu, Peng Zhang, Bang Zhang, Pan Pan et al.CVPR 2021
- Per-Clip Video Object SegmentationKwanyong Park, Sanghyun Woo, Seoung Wug Oh, In So Kweon et al.CVPR 2022 · 45 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Alignment Before Aggregation: Trajectory Memory Retrieval Network for Video Object SegmentationRui Sun, Yuan Wang, Huayu Mai, Tianzhu Zhang et al.ICCV 2023 · 12 citations
- Space-Time Distillation for Video Super-ResolutionZeyu Xiao, Xueyang Fu, Jie Huang, Zhen Cheng et al.CVPR 2021
